←  Back to all vacancies

Senior Data Engineer

Mirai Arabian International

Media & Entertainment

📍 Saudi Arabia
💼 Full-time
🕒 Posted 3 months ago

Job description

Role Overview

Senior Data Engineer at Mirai Arabian International. This role owns the complete data layer for generative AI products: the pipelines that bring data in, the transformations that shape it, and the way it reaches retrieval systems, agents, and analytics. The work runs on AWS.

Role Purpose

Build and operate a governed, single source of truth for data across the organization. Prepare data specifically for AI systems—including chunking, embeddings, and indexing—not only for reporting. Lead the team's data engineering standards and practices.

Key Responsibilities

Pipelines & Infrastructure

  • Build and run batch and streaming pipelines that move data from source systems into the lake and through to the warehouse.
  • Own the layers in between from raw to curated, including schema, quality, and lineage.
  • Build the data layer behind retrieval: source connectors, document parsing, chunking, embedding generation, and vector indexing.
  • Re-embed content when it changes.
  • Work with platform and DevOps engineers to expose data and retrieval as documented, dependable services.

Data Modeling & Curation

  • Model curated, query-ready datasets and metrics so AI and analytics consumers work from one definition.
  • Prevent each consumer from rebuilding the same logic.

Quality & Governance

  • Add quality checks, validation, and monitoring so problems surface before they reach a model or user.
  • Apply access control: row and column level rules, PII handling, and entitlement-aware datasets.
  • Enforce access control as close to query time as the stack allows.

Cost & Performance

  • Keep storage, compute, and query costs in check.
  • Pay particular attention to the cost of embedding and vector workloads.

Team & Standards

  • Review code, write documentation, and help shape how the team builds its data layer.

Qualifications & Experience

  • Eight or more years in data engineering overall.
  • Hands-on work building data for AI or ML systems such as retrieval, embeddings, or feature data (can be a more recent part of your background).
  • Production experience across the AWS data stack: S3 for the lake, Glue for ETL and the Data Catalog, Athena for serverless query, and Redshift as the warehouse.
  • Hands-on experience with a layered data architecture, whether medallion (bronze, silver, gold), a data lake feeding a warehouse, or a lakehouse.
  • Experience building the transformation stages that move data from raw to curated.
  • Experience with an ELT or integration tool such as Airbyte, Fivetran, or Meltano, including building or maintaining connectors.
  • Experience with event-driven pipelines using SQS and SNS.
  • Experience with at least one streaming or change-data-capture technology such as Kinesis, Amazon MSK, or Debezium.
  • Hands-on experience with a semantic or metrics layer over the warehouse, such as Cube or the dbt Semantic Layer.
  • Hands-on experience with at least one vector store and embedding workflow: pgvector, Amazon OpenSearch, Pinecone, Weaviate, or Milvus.
  • Working knowledge of columnar and open table formats: Parquet together with Apache Iceberg, Delta Lake, or Hudi.

Skills & Competencies

  • Strong SQL and strong Python.
  • PySpark or similar distributed processing.
  • Working knowledge of an orchestrator such as Amazon MWAA, Step Functions, Dagster, or Prefect.
  • Enough infrastructure as code to work closely with DevOps.

People looking at this role also searched

Report this job

⚡ Quick Apply

Create your account and upload your CV to apply for — takes less than a minute.

✨ Get a free AI ATS Score Report for your CV the moment you sign up.