وصف الوظيفة
Role Overview
Associate AI Quality Engineer at Foodics, a leading restaurant management ecosystem and payment tech provider headquartered in Riyadh.
Company Overview
Foodics was founded in 2014 and operates offices across 5 countries, including UAE, Egypt, Jordan, and Kuwait, serving customers and partners in over 35 countries worldwide. The company has processed over 6 billion orders and is one of the most rapidly evolving SaaS companies from the MENA region. Foodics has completed three funding rounds, with the latest raising $170 million in the largest SaaS funding round in MENA.
Role Purpose
You will build internal AI systems that enable engineers to work more effectively every day: agents that generate and maintain tests, pipelines that triage failures before human review, tooling that accelerates code review and debugging, and evaluation infrastructure that makes internal AI features testable. You will write production code, own systems in CI, and be measured on whether engineers actually use what you ship.
Key Responsibilities
Test Automation
- Build API and contract testing across the backend stack with service-level and integration coverage.
- Design test data setup that remains maintainable and test design that survives schema changes.
- Develop web end-to-end and component-level coverage using Playwright with visual and RTL regression testing.
- Create frontend test suites fast enough to gate merges rather than run nightly.
- Implement native and cross-platform mobile coverage using Appium or Maestro with device-farm strategy.
- Test offline and sync behaviour and payment-peripheral paths that only surface on real hardware.
- Establish shared fixtures, environment and test-data management across the automation stack.
- Implement parallelization and CI pipelines where a red build carries meaningful signal.
- Deliver performance and load testing experience.
AI-Driven Quality Systems
- Generate tests from specifications, code, and production traffic with solved maintenance stories, not just initial drafts.
- Build failure triage that classifies red builds before human review: real bug, flake, environment issue, or test rot.
- Develop self-healing locators and suite health tooling including flake detection, quarantine, and coverage-gap analysis.
- Establish evaluation infrastructure for AI features across products: datasets, scoring, and regression detection when prompts or models change.
- Evaluate performance for the market: Arabic and English behaviour, RTL interfaces, and region-specific POS, tax, and payment rules where correctness is rarely a string match.
Agentic AI and Orchestration
- Build agentic AI that executes real work in pipelines: reads diffs, runs relevant suites, reproduces failures, proposes fixes, and opens pull requests.
- Own agents that manage quality workflows end to end—exploratory testing against running builds, coverage-gap hunting, release-risk assessment—and know when to escalate to humans.
- Implement orchestration that sustains load: multi-step planning, tool use, retries, state and memory across steps, sandboxed execution, multi-agent handoffs, and clean boundaries between agentic and deterministic steps.
- Integrate with existing stack (CI, Jira, observability, MCP-style tool interfaces) rather than building parallel systems.
- Exercise judgement to know when a plain pipeline outperforms an agent and communicate that assessment.
Technical Foundation
- Operate with current knowledge of test automation: framework design and layering, test pyramid application, flake economics, parallel execution, mobile and cross-browser realities, CI/CD gating.
- Apply expertise in agentic AI: orchestration and tool use, multi-step planning, memory and state, sandboxed execution, multi-agent patterns, MCP and similar tool-integration standards, and cost analysis of each.
- Manage context engineering: retrieval strategy, chunking, reranking, caching, and long-context behaviour including degradation points.
- Execute evaluation: offline and online evals, LLM-as-judge and its failure modes, human-in-the-loop review, statistical significance on small samples, regression gates in CI.
- Ensure reliability: structured output, guardrails, fallback and retry design, and handling non-determinism in systems that cannot flap.
- Operate LLM systems: tracing and observability, prompt and version management, latency and cost budgeting, model routing, and knowing when fine-tuning or distillation beats improved prompting.
Qualifications & Experience
- Recent hands-on work on LLM-backed systems that real users depend on.
- Real automation depth across more than one surface; owned a suite that gates releases on backend and on a UI (web or mobile) and kept it green without deleting hard tests.
- Real experience building evaluation systems with demonstrated quantified improvement.
- Practical depth with the modern LLM toolkit: prompting, structured output, tool use, retrieval, and agentic AI orchestration with clear understanding of trade-offs.
- Credible testing fundamentals; test design, automation frameworks, and CI/CD should not be new territory.
Skills & Competencies
- Strong Python proficiency.
- Comfortable in at least one of .NET, Java, or TypeScript.
- Tested, maintained code experience; not notebook-based work.
- Hands-on expertise with Playwright, Appium, and Maestro frameworks.
- Bias toward adoption; measure work by actual engineer usage, not by demonstrations.
Additional Information
- Hybrid work setup with flexibility.
- Highly competitive compensation packages including bonuses and potential shares.
- Regular training and annual learning stipend.
- Autonomy, mentoring, and challenging goals.
- Diverse team of over 30 nationalities working across 14 countries.