Data Engineer (43602)
Become part of AI incubator projects focused on exploring and validating new client and internal ideas and creating advanced AI products, services, and capabilities. As a Data Engineer, you will design and operate production-ready data pipelines and support the full Data Science and ML lifecycle. You will need strong Data Engineering and ETL/ELT experience, hands-on Python, PySpark, Databricks, and Azure skills, and experience with data quality, validation, monitoring, and production readiness.
🚀 Project
- participate in AI incubator projects to scout, incubate, and validate client and internal ideas on a 3–5-year horizon
- develop technology roadmaps and prototypes that deliver advanced client solutions
- champion internally generated concepts
- continually explore, test, and demonstrate cutting-edge AI to create new products, services, and capabilities
- design, build, and operate reliable, production-ready data pipelines for large-scale data and ML use cases
- bridge raw data ingestion, transformation, quality assurance, and downstream machine learning workflows
- support the end-to-end Data Science and ML lifecycle from raw ingestion and hardening through production deployment, monitoring, and continuous improvement
- work with NLP/NLU use cases, Azure data and AI services, Databricks/PySpark processing, CI/CD, and MLOps practices
- build, harden, and maintain reliable ingestion pipelines for production-ready data platforms
- design and implement ETL/ELT processes for data cleaning, transformation, schema design, data quality validation, and lineage tracking
- develop large-scale data processing and transformation workflows using Databricks and PySpark
- prepare high-quality datasets, features, and pipelines for NLP, NLU, and broader machine learning use cases
- operate and integrate Azure data and AI infrastructure, including Blob Storage, databases, compute resources, MLflow, and Azure AI/ML services
- implement CI/CD and MLOps practices including automated deployment, testing, monitoring, reliability checks, and promotion gates
- ensure pipeline reliability, observability, performance, and production readiness across the full data and ML lifecycle
- own data quality, traceability, and operational handover for production data products and ML workflows
🎯 Skills
- strong Data Engineering experience including ETL/ELT, pipeline development, data cleaning, transformation, schema design, and orchestration
- hands-on experience building reliable, production-ready data ingestion pipelines with hardening, validation, monitoring, and operational controls
- strong Python and PySpark skills for large-scale data processing, transformation, and pipeline development
- practical experience with Databricks for scalable data engineering, distributed processing, and production pipeline implementation
- hands-on Azure infrastructure experience including Blob Storage, databases, compute resources, MLflow, and Azure AI/ML services
- solid understanding of data quality, lineage tracking, observability, reliability, and production readiness for data pipelines
- strong Data Science and ML lifecycle awareness covering model development, experimentation, deployment, monitoring, and continuous improvement
- experience supporting NLP and NLU use cases, including preparation of high-quality datasets and pipelines for natural language applications
- English at C1 level
💡 Nice to have
- experience implementing CI/CD and MLOps practices including automated testing, deployment, monitoring, promotion gates, and release reliability
- ability to bridge Data Engineering and Data Science teams and translate ML workflow needs into reliable data products and production pipelines
- experience with enterprise-grade ML platforms, feature pipelines, experiment tracking, model governance, or production ML operations
- familiarity with regulated or complex enterprise environments where data quality, traceability, security, and operational resilience are critical
#LI-MP9