←  Back to all vacancies

Lead Data Intelligence Machine Learning Engineer

Dyson

Retail & FMCG

πŸ“ UAE
πŸ’Ό Full-time
πŸ•’ Posted 3 weeks ago

Job description

Role Overview

Lead Data Intelligence Machine Learning Engineer at Dyson. This is a specialist position focused on designing and implementing in-house tools that automate data labelling pipelines to reduce reliance on manual annotation and ensure high-quality labels at scale.

Company Overview

Dyson is driven by a relentless pursuit of innovation in engineering, AI, and robotics. The new Data Intelligence team sits at the heart of this mission, shaping Dyson's future through data by blending creativity, precision, and audacity to power intelligent products. You will work alongside brilliant minds from Dyson's global engineering team and external software and hardware partners in an environment built for exploration, discovery, delivery, and impact.

Role Purpose

To architect, develop, and maintain automated data labelling systems that bridge the gap between raw data collection and model-ready datasets. This role leverages techniques such as Active Learning, Weak Supervision, and Synthetic Data Generation to scale high-quality data annotation while reducing manual effort and ensuring seamless integration with downstream MLOps infrastructure.

Key Responsibilities

Pipeline Architecture & Development

  • Architect labelling pipelines by designing and deploying end-to-end automated labelling systems using frameworks like Snorkel, Cleanlab, or custom active learning loops.
  • Develop human-in-the-loop (HITL) systems that build interfaces and workflows where models pre-label data and humans intervene only on high-uncertainty samples.
  • Set up and maintain robust data preparation infrastructure optimised for data quality, speed, and seamless integration with downstream MLOps pipelines.

Quality Assurance & Data Integrity

  • Implement algorithmic checks to identify and correct mislabelled or noisy data within existing datasets.
  • Fine-tune teacher models to generate high-quality pseudo-labels for student models.

Integration & Infrastructure

  • Collaborate with software engineers to integrate labelling tools with existing data lakes and ML training infrastructure.
  • Design, deploy, and maintain scalable data pipelines, including data cleansing, transformation, and storage on cloud, on-premises, or hybrid environments.

Analysis & Insight

  • Perform data visualization and in-depth analysis using advanced data and feature engineering techniques.
  • Transform raw data into actionable insight, supporting both research and deployment.

Stakeholder Collaboration

  • Work closely with Data Scientists, Software Engineers, and Product teams to ensure high data quality and usability across products and projects.
  • Communicate solutions clearly and collaborate effortlessly across technical and non-technical teams.

Qualifications & Experience

  • At least 8 years of professional experience in Machine Learning engineering, specifically focused on data-centric AI or computer vision and NLP pipelines.
  • Bachelor's or Master's degree in Computer Science, Engineering, Mathematics, Data Science, or a related field.
  • Proven hands-on expertise building auto-labelling solutions or working with large-scale data annotation workflows.
  • Experience with Weak Supervision (labelling functions) or Active Learning strategies (uncertainty sampling, diversity sampling).
  • Experience with SQL and NoSQL databases and managing large-scale unstructured data (images, text, or audio).
  • Experience designing, deploying, and maintaining scalable data pipelines, including data cleansing, transformation, and storage.
  • Familiarity with AWS (SageMaker Ground Truth), GCP (Vertex AI), or Azure ML labelling services.
  • Experience with DVC (Data Version Control) or similar tools to track dataset iterations.

Skills & Competencies

  • Mastery of Python and the Machine Learning stack: PyTorch or TensorFlow, NumPy, Pandas, and Scikit-learn.
  • Advanced proficiency with key ML and data science libraries including TensorFlow, PyTorch, and scikit-learn.
  • Strong background in feature engineering, data analysis, and data visualization using tools such as Jupyter, Tableau, or Power BI.
  • Expertise in Weak Supervision and Active Learning methodologies.
  • Ability to balance speed and quality in a fast-moving environment.
  • Stay curious about new developments and deliver results under pressure.
  • Strong communicator who documents solutions clearly.

Additional Information

Dyson is an equal opportunity employer. Employment decisions are made without regard to race, colour, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status, or any other dimension of diversity.

People looking at this role also searched

Report this job

⚑ Quick Apply

Create your account and upload your CV to apply for β€” takes less than a minute.

✨ Get a free AI ATS Score Report for your CV the moment you sign up.