Job description
Role Overview
Salla is hiring a Senior SRE Engineer (MLOps) to join the Salla AI team. This is a production infrastructure and reliability engineering role focused on AI and ML systems.
Company Overview
Salla is building Agentic AI and Generative AI features within the Salla ecosystem. The AI team operates production AI systems, models, agents, and inference services at scale.
Role Purpose
This role exists to run AI and ML systems as real production systems—not experiments. You will own the operational layer around models, prompts, agents, inference services, and retrieval systems, ensuring they operate reliably, securely, and cost-effectively. AI systems fail differently from normal services; a prompt change behaves like a code change, agents need auditability, and latency, quality, and cost move together in uncomfortable ways. Your work gives product teams a fast, safe path to production.
Key Responsibilities
Reliability & Observability
- Own reliability for ML and agentic AI services in production, including SLOs, dashboards, alerts, runbooks, and incident follow-ups.
- Build observability across the AI stack, covering latency, errors, traces, tool calls, cost, and user impact.
- Provide operational support for inference APIs, queues, retrieval layers, and AI workflows running on Kubernetes/EKS.
Safe Releases & Deployment
- Design safe-release patterns for models, prompts, agents, tools, and configuration, including canary, rollback, feature-flag, and evaluation-gate strategies.
- Build automation and self-service paths so product teams have a known safe path to production instead of rebuilding it each time.
Security & Governance
- Establish ownership, traceability, and guardrails around what agentic systems are allowed to do, including how they call internal tools.
- Defend agent tool-calling against prompt injection and untrusted-data risks.
- Establish and enforce data-trust boundaries so that untrusted store or merchant content cannot manipulate agent decisions, tool calls, or actions.
- Drive AI cost governance through per-model and per-pod spend visibility, token-cost tracking, and anomaly alerting.
Platform & Standards
- Turn recurring operational pain into simple, reusable platform standards that other teams adopt.
- Participate in architecture discussions, code reviews, and technical decision-making.
Qualifications & Experience
- 4+ years in SRE, platform engineering, DevOps, or production infrastructure operating distributed systems in production—not only in demos.
- Hands-on experience with Kubernetes and cloud-native systems in production.
- Familiarity with deploying ML projects.
- Strong command of CI/CD, GitOps, observability, and incident response.
- Solid experience with infrastructure-as-code, secrets management, and networking.
- Production judgment—knowing how to make systems measurable, debuggable, repeatable, and safe to change. You do not need to be a machine learning researcher.
Skills & Competencies
- Ability to write automation or platform tooling in Python or a similar language.
- Ability to work across teams, explain trade-offs clearly, and turn operational pain into standards engineers will actually use.
Nice to Have
- Experience with MLOps or ML platforms—model serving, registries, evaluation, feature or data dependencies, drift monitoring, or ML pipelines.
- Familiarity with LLM applications or agentic systems—RAG, vector databases, tool calling, workflow orchestration, memory, traces, guardrails, or evaluation pipelines.
- Exposure to tooling such as OpenTelemetry, Prometheus, Grafana, MLflow, KServe, Ray, LiteLLM, vLLM, LangGraph, Arize Phoenix, or LangSmith.
- Experience with Kafka consumers, GPU workloads, inference optimization, model routing, or AI cost governance.
- Experience working in cross-functional product teams involving AI, backend, and frontend engineers.