Job description
Role Overview
Senior AI Backend Engineer - Agent Evaluation & Quality at Salla. This role owns the evaluation systems for production multi-agent systems, building the judges, test harnesses, and simulators that measure agent performance and catch regressions before release.
Role Purpose
You will design and operate the evaluation stack that enables confident, fast agent releases. Your systems will measure agent quality, enforce quality gates in CI/CD, catch regressions automatically, and compound improvement over time by turning production failures into better test cases. As you deepen your understanding of agent failure modes, you will also contribute directly to agent development and hardening.
Key Responsibilities
Evaluation Infrastructure
- Own the evaluation stack end-to-end.
- Design and build LLM-as-judge systems calibrated against human labels.
- Make agent quality measurable per-agent and per-failure-mode.
- Build per-PR eval harnesses and regression detection wired into CI so quality is enforced automatically.
Testing & Coverage
- Build user simulators to generate test coverage and adversarial cases before real users encounter them.
- Turn production signal into improvement by piping real failures back into evaluation sets so the system compounds over time.
Stakeholder Alignment
- Partner with product to turn "what good looks like" into concrete, measurable criteria.
Agent Development
- Grow into agent development by contributing to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.
Qualifications & Experience
- 5+ years of software engineering with recent hands-on LLM/agent work.
- Production experience running LLM systems and dealing with reliability, latency, cost, and observability.
Skills & Competencies
- Strong software engineering fundamentals: production Python or Typescript (or similar), clean API and system design, testing, and CI/CD.
- Hands-on LLM/agent experience: you have built with LLMs including agents, RAG, tool/function calling, and orchestration frameworks (LangGraph, LangChain, or equivalent), and understand how they behave and break.
- A measurement mindset: you reason about metrics, calibration, and experiments, and want to quantify whether something works rather than ship untested.
- You write code others build on; evaluation infrastructure is real engineering.
Additional Information
Nice to Have
- Direct experience evaluating LLM/agent systems, including offline/online eval, LLM-as-judge, and systematic regression testing.
- Observability tooling experience (Arize, LangSmith, or similar).
- Arabic language or NLP experience.
- E-commerce or merchant-facing product experience.