Job description
Role Overview
Data Intelligence Machine Learning Engineer at Dyson. The role is based within Dyson's new Data Intelligence team, which sits at the heart of the company's mission to shape Dyson's future through data.
Company Overview
Dyson is driven by a relentless pursuit of innovation, pushing boundaries in engineering, AI, and robotics. The Data Intelligence team blends creativity, precision, and audacity to power intelligent products, crafting data strategies and pipelines that fuel the next generation of connected devices. You will work alongside brilliant minds from Dyson's global engineering team and external software and hardware partners in an environment built for exploration, discovery, delivery, and impact.
Role Purpose
The primary goal of this role is to design and implement in-house tools that automate Dyson's data labelling pipelines, reducing reliance on manual annotation by leveraging techniques such as Active Learning, Weak Supervision, and Synthetic Data Generation. You will bridge the gap between raw data collection and model-ready datasets, ensuring high-quality labels at scale.
Key Responsibilities
Labelling Pipeline Architecture
- Architect end-to-end automated labelling systems using frameworks like Snorkel, Cleanlab, or custom active learning loops.
- Design and deploy labelling pipelines that scale across the organization.
Human-in-the-Loop Systems
- Develop Human-in-the-Loop (HITL) systems and workflows where models pre-label data and humans only intervene on high-uncertainty samples.
- Build interfaces and tools to support HITL workflows.
Data Quality & Denoising
- Implement algorithmic checks to identify and correct mislabelled or noisy data within existing datasets.
- Perform quality assurance across data preparation processes.
Infrastructure & Integration
- Set up and maintain robust data preparation infrastructure, optimizing for data quality, speed, and seamless integration with downstream MLOps pipelines.
- Collaborate with software engineers to integrate labelling tools with existing data lakes and ML training infrastructure.
Model Optimization
- Fine-tune teacher models to generate high-quality pseudo-labels for student models.
Data Analysis & Visualization
- Perform data visualization and in-depth analysis using advanced data and feature engineering techniques.
- Transform raw data into actionable insight, supporting both research and deployment.
Cross-functional Collaboration
- Work closely with Data Scientists, Software Engineers, and Product teams to ensure high data quality and usability across products and projects.
Qualifications & Experience
- At least 3 years of professional experience in Machine Learning engineering, specifically focused on data-centric AI or computer vision and NLP pipelines.
- Bachelor's or Master's degree in Computer Science, Engineering, Mathematics, Data Science, or a related field.
- Proven hands-on expertise building auto-labelling solutions or working with large-scale data annotation workflows.
- Experience designing, deploying, and maintaining scalable data pipelines, including data cleansing, transformation, and storage in cloud, on-premises, or hybrid environments.
Skills & Competencies
Programming & Machine Learning
- Advanced proficiency in Python with mastery of the Machine Learning stack: PyTorch or TensorFlow, NumPy, Pandas, and Scikit-learn.
- Experience with other relevant programming languages as applicable.
Automated Labelling & Active Learning
- Proven experience with Weak Supervision, including labelling functions.
- Proven experience with Active Learning strategies, including uncertainty sampling and diversity sampling.
Data Engineering
- Experience with SQL and NoSQL databases.
- Ability to manage large-scale unstructured data, including images, text, and audio.
- Experience with DVC (Data Version Control) or similar tools to track dataset iterations.
Cloud Infrastructure
- Familiarity with AWS (SageMaker Ground Truth), GCP (Vertex AI), or Azure ML labelling services.
Data Analysis & Visualization
- Strong background in feature engineering, data analysis, and data visualization.
- Comfortable using tools such as Jupyter, Tableau, or Power BI.
Soft Skills
- Excellent communication skills with ability to document solutions clearly.
- Ability to collaborate effortlessly across technical and non-technical teams.
- Ability to balance speed and quality in a fast-moving environment.
- Demonstrated curiosity about new developments in the field.
Additional Information
Dyson is an equal opportunity employer. The company welcomes applications from all backgrounds and makes employment decisions without regard to race, colour, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status, or any other dimension of diversity.