←  Back to all vacancies

Senior AI Backend Engineer - Agent Evaluation & Quality

Salla

Technology & IT

πŸ“ Saudi Arabia
πŸ’Ό Full-time
πŸ•’ Posted 2 months ago

Job description

Role Overview

Senior AI Backend Engineer - Agent Evaluation & Quality at Salla. This role owns the evaluation systems for production multi-agent systems, building the judges, test harnesses, and simulators that measure agent performance and catch regressions before release.

Role Purpose

You will design and operate the evaluation stack that enables confident, fast agent releases. Your systems will measure agent quality, enforce quality gates in CI/CD, catch regressions automatically, and compound improvement over time by turning production failures into better test cases. As you deepen your understanding of agent failure modes, you will also contribute directly to agent development and hardening.

Key Responsibilities

Evaluation Infrastructure

  • Own the evaluation stack end-to-end.
  • Design and build LLM-as-judge systems calibrated against human labels.
  • Make agent quality measurable per-agent and per-failure-mode.
  • Build per-PR eval harnesses and regression detection wired into CI so quality is enforced automatically.

Testing & Coverage

  • Build user simulators to generate test coverage and adversarial cases before real users encounter them.
  • Turn production signal into improvement by piping real failures back into evaluation sets so the system compounds over time.

Stakeholder Alignment

  • Partner with product to turn "what good looks like" into concrete, measurable criteria.

Agent Development

  • Grow into agent development by contributing to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.

Qualifications & Experience

  • 5+ years of software engineering with recent hands-on LLM/agent work.
  • Production experience running LLM systems and dealing with reliability, latency, cost, and observability.

Skills & Competencies

  • Strong software engineering fundamentals: production Python or Typescript (or similar), clean API and system design, testing, and CI/CD.
  • Hands-on LLM/agent experience: you have built with LLMs including agents, RAG, tool/function calling, and orchestration frameworks (LangGraph, LangChain, or equivalent), and understand how they behave and break.
  • A measurement mindset: you reason about metrics, calibration, and experiments, and want to quantify whether something works rather than ship untested.
  • You write code others build on; evaluation infrastructure is real engineering.

Additional Information

Nice to Have

  • Direct experience evaluating LLM/agent systems, including offline/online eval, LLM-as-judge, and systematic regression testing.
  • Observability tooling experience (Arize, LangSmith, or similar).
  • Arabic language or NLP experience.
  • E-commerce or merchant-facing product experience.

People looking at this role also searched

Report this job

⚑ Quick Apply

Create your account and upload your CV to apply for β€” takes less than a minute.

✨ Get a free AI ATS Score Report for your CV the moment you sign up.