وصف الوظيفة
Role Overview
Senior Data Intelligence Machine Learning Engineer at Dyson, focused on designing and implementing automated data labelling pipelines. The role is based within Dyson's Data Intelligence team, which sits at the core of the company's innovation mission across engineering, AI, and robotics.
Company Overview
Dyson is driven by a relentless pursuit of innovation in engineering, AI, and robotics. The Data Intelligence team shapes Dyson's future through data, blending creativity, precision, and audacity to power intelligent products. The team crafts data strategies and pipelines that fuel the next generation of connected devices and works alongside Dyson's global engineering team and external software and hardware partners in an environment built for exploration, discovery, delivery, and impact.
Role Purpose
To design and implement in-house tools that automate data labelling pipelines, reducing reliance on manual annotation by leveraging techniques such as Active Learning, Weak Supervision, and Synthetic Data Generation. The role bridges the gap between raw data collection and model-ready datasets, ensuring high-quality labels at scale.
Key Responsibilities
Architecture & Pipeline Design
- Architect labelling pipelines: Design and deploy end-to-end automated labelling systems using frameworks like Snorkel, Cleanlab, or custom active learning loops.
- Set up and maintain robust data preparation infrastructure, optimizing for data quality, speed, and seamless integration with downstream MLOps pipelines.
- Experience designing, deploying, and maintaining scalable data pipelines, including data cleansing, transformation, and storage in cloud, on-premise, or hybrid environments.
Human-in-the-Loop Systems
- Develop "Human-in-the-Loop" (HITL) systems: Build interfaces and workflows where models pre-label data and humans only intervene on high-uncertainty samples.
Quality Assurance & Data Denoising
- Implement algorithmic checks to identify and correct mislabelled or "noisy" data within existing datasets.
Integration & Tooling
- Collaborate with software engineers to integrate labelling tools with existing data lakes and ML training infrastructure.
Model Optimization
- Fine-tune "teacher" models to generate high-quality pseudo-labels for "student" models.
Data Analysis & Visualization
- Perform data visualization and in-depth analysis using advanced data and feature engineering techniques.
- Transform raw data into actionable insight, supporting both research and deployment.
- Develop strong background in feature engineering, data analysis, and data visualization, comfortable using tools like Jupyter, Tableau, or Power BI.
Collaboration & Communication
- Work closely with Data Scientists, Software Engineers, and Product teams to ensure high data quality and usability across products and projects.
- Communicate clearly and document solutions, collaborating effortlessly across technical and non-technical teams.
Qualifications & Experience
- Bachelor's or Master's degree in Computer Science, Engineering, Mathematics, Data Science, or a related field.
- At least 5+ years of professional experience in Machine Learning engineering, specifically focused on data-centric AI or computer vision and NLP pipelines.
- Proven experience with Weak Supervision (labelling functions) or Active Learning strategies (uncertainty sampling, diversity sampling).
- Hands-on expertise building auto-labelling solutions or working with large-scale data annotation workflows.
- Experience with SQL and NoSQL databases, and managing large-scale unstructured data (images, text, or audio).
- Familiarity with AWS (SageMaker Ground Truth), GCP (Vertex AI), or Azure ML labelling services.
- Experience with DVC (Data Version Control) or similar tools to track dataset iterations.
Skills & Competencies
- Proficiency in Python: Mastery of the Machine Learning stack (PyTorch or TensorFlow, NumPy, Pandas, Scikit-learn).
- Advanced skills in Python and other relevant languages, with experience in key ML and data science libraries such as TensorFlow, PyTorch, scikit-learn, and pandas.
- Strong ability to balance speed and quality, stay curious about new developments, and deliver results in a fast-moving environment.
Additional Information
- Dyson is an equal opportunity employer. Employment decisions are made without regard to race, colour, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status, or any other dimension of diversity.