وصف الوظيفة
Role Overview
Data Engineer at Deeplight, a specialist AI and data consultancy. Based in the UAE, this role joins the core engineering squad to design and optimize scalable data solutions for enterprise analytics and machine learning at production scale.
Company Overview
DeepLight AI is a specialist AI and data consultancy dedicated to transforming the regional corporate landscape through bespoke, high-impact intelligent systems. The firm combines deep expertise in data engineering, cloud architecture, AI/ML platforms, and systems integration with practical understanding of complex enterprise operations. With deep roots in financial services and banking sectors, Deeplight helps major institutions bridge the gap between complex data infrastructure and actionable AI strategy, engineering secure, scalable, and resilient data foundations designed to power production AI at scale.
Role Purpose
You will design, construct, and optimize scalable Lakehouse data solutions that power enterprise analytics and advanced machine learning models. Operating at the intersection of big data, cloud architecture, and MLOps, you will unify data warehousing and data lake capabilities to process complex, high-volume financial data while maintaining strict banking security, ACID compliance, and data governance standards.
Key Responsibilities
Data Architecture & Lakehouse Design
- Build next-generation Lakehouse architectures using Databricks and Delta Lake to unify data warehousing and data lake capabilities on AWS.
- Design and construct scalable data solutions optimized for enterprise-grade performance.
Pipeline Development
- Develop high-throughput batch and real-time ETL/ELT pipelines using PySpark, Apache Kafka, and SQL to ingest and transform complex financial datasets from core banking, market data, and transactional systems.
Performance & Cost Optimization
- Optimize query performance and storage efficiency across AWS services (S3, Redshift, EMR, Athena, Glue) using advanced partitioning, indexing, and parallel processing to drive down latency and cloud compute costs.
Machine Learning & MLOps Integration
- Partner with Data Scientists and MLOps Engineers to build automated feature engineering pipelines, track experiments using MLflow, containerize workflows with Docker and EKS, and deploy models to production using AWS SageMaker.
Data Governance & Compliance
- Implement end-to-end data security, encryption, access controls, and audit trails to ensure 100% compliance with Central Bank data sovereignty and privacy mandates.
Engineering Practices & Automation
- Drive software engineering best practices by managing Infrastructure as Code with Terraform, authoring CI/CD pipelines using GitHub Actions or CodePipeline, and writing clean, well-tested Python code.
Stakeholder Communication
- Bridge the gap between complex data science and actionable business value through compelling communication and persuasive advocacy for technical decisions.
- Articulate complex technical concepts clearly to cross-functional teams and client stakeholders.
- Present cutting-edge solutions and build trust with high-level stakeholders.
Qualifications & Experience
- Hands-on data engineering experience directly within banking, financial services, or fintech environments, with strong understanding of transactional datasets and regulatory compliance.
- Proven track record building and managing Lakehouse architectures with Databricks or Delta Lake, including expert knowledge of data modeling, schema evolution, and ACID principles.
Skills & Competencies
- Deep technical command of core AWS data services (S3, Glue, EMR, Redshift).
- Expert proficiency with distributed processing engines (Apache Spark, PySpark, Apache Kafka).
- High proficiency in Python and SQL; Scala is a plus.
- Practical experience with Docker, Kubernetes (EKS), Terraform, and Git workflows.
- Excellent analytical problem-solving skills.
- Ability to bridge complex data science concepts with actionable business value.
- Strong communication skills and ability to articulate technical decisions to stakeholders.
Preferred Skills
- AWS Certified Data Engineer, AWS Solutions Architect, or Databricks Certified Data Engineer certification.
- Hands-on experience with streaming analytics frameworks (Apache Flink, Spark Streaming) for real-time fraud detection or trade processing pipelines.
- Production experience with modern workflow schedulers such as Apache Airflow or Prefect.
Additional Information
- Competitive salary.
- Comprehensive 'Gold Level' personal health insurance.
- Individual Visa Sponsorship.
- Professional development and certification support.
- Subscription reimbursement relating to your role.
- Opportunity to work on cutting-edge AI projects.
- Monthly Employee Incentive program.
- Career advancement opportunities in a rapidly growing AI company.