←  Back to all vacancies

Site Reliability Engineer

Lucidya

Technology & IT

πŸ“ Saudi Arabia
πŸ’Ό Full-time
πŸ•’ Posted 3 weeks ago

Job description

Role Overview

Site Reliability Engineer at Lucidya. Lucidya is an AI-native platform for customer experience (CX) intelligence that manages entire customer lifecycles autonomously, from initial engagement through retention and growth.

Company Overview

Lucidya is an AI-native platform for customer experience (CX) intelligence that processes massive volumes of real-time customer data. The platform uses proprietary NLU technology built in-house and trained on millions of multilingual conversations, enabling marketing, support, CX, and research teams to deliver personalized experiences that drive measurable improvements in customer satisfaction, retention, and lifetime value. As the company scales globally, the reliability, performance, and resilience of infrastructure become mission-critical.

Role Purpose

As a Site Reliability Engineer, you will own the reliability of Lucidya's cloud infrastructure and ensure it scales seamlessly as the organization grows. You will anticipate infrastructure challenges, design systems that prevent failures, and build automation that removes issues entirely. Any downtime, latency, or instability directly impacts customers' ability to make decisions and serve their own users; this role ensures that does not happen.

Key Responsibilities

Infrastructure Design & Reliability

  • Design and maintain infrastructure that is highly available, fault-tolerant, and scalable.
  • Proactively identify and eliminate single points of failure before they become incidents.
  • Ensure production systems remain stable, even under increasing scale and load.

Cloud Environment Management

  • Manage and continuously improve workloads across AWS, GCP, or Azure.
  • Use Infrastructure as Code (Terraform) to standardize and scale infrastructure.
  • Optimize resource usage to balance performance and cost.

Kubernetes Operations

  • Operate and scale Kubernetes clusters (EKS, GKE, etc.) with confidence.
  • Troubleshoot issues quickly and ensure smooth deployments and upgrades.
  • Ensure containerized workloads perform reliably at scale.

Observability & Incident Management

  • Implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK.
  • Define alerting that is meaningful, not noisy.
  • Respond to incidents, lead root cause analysis, and ensure the organization learns from every failure.

Automation & Efficiency

  • Write scripts and build tooling to eliminate repetitive operational work.
  • Continuously improve infrastructure efficiency through automation.
  • Promote a culture where manual work is a temporary state, not the norm.

Collaboration & Best Practices

  • Work closely with DevOps and engineering teams to solve performance bottlenecks.
  • Contribute to CI/CD improvements and deployment reliability.
  • Help shape reliability best practices across the organization.

Qualifications & Experience

  • Approximately 3 years of hands-on experience working in SRE, DevOps, or infrastructure engineering, with exposure to systems at scale.
  • Hands-on experience working with Kubernetes in production and troubleshooting when issues arise.
  • Comfortable working in cloud environments like AWS, GCP, or Azure with understanding of how distributed systems behave.

Skills & Competencies

Infrastructure & Cloud Tools

  • Use Terraform or similar Infrastructure as Code tools to manage infrastructure.
  • Work confidently with Docker and Kubernetes.
  • Understand networking, load balancing, and high-availability design.
  • Understand CI/CD pipelines (Jenkins, GitHub Actions, Bitbucket, etc.).

Scripting & Automation

  • Write scripts in Python, Bash, or similar languages to automate workflows.

Monitoring & Observability

  • Implement tools like Prometheus, Grafana, Datadog, or ELK.
  • Know the difference between useful alerts and noise.
  • Focus on signals that actually drive action.

Personal Qualities

  • Take ownershipβ€”do not wait to be told something is broken.
  • Remain calm under pressure and methodical during incidents.
  • Simplify complexity instead of adding to it.
  • Communicate clearly, even when explaining deeply technical issues.
  • Care about building systems that make other engineers more effective.

Additional Information

Nice to Have

  • Experience with RabbitMQ or Redis in production.
  • Familiarity with Ansible or AWX.
  • Exposure to multi-cloud or hybrid environments.
  • Cloud certifications (AWS, GCP) or Linux certifications.
  • Background from ITI (Information Technology Institute).

First 90 Days

By day 30, you will have built a strong understanding of Lucidya's infrastructure, systems, and workflows; contribute to day-to-day operations with support from the team; and start identifying areas for improvement in automation and reliability.

By day 90, you will independently manage infrastructure tasks and troubleshoot issues, actively contribute to reliability and scalability improvements, and take ownership of parts of the infrastructure while improving them.

Hiring Process

  • Screening Interview – Talent Acquisition.
  • Technical Interview – SRE Lead.
  • Technical Task.
  • Final Interview – SRE Lead & Cloud DevOps Director.

People looking at this role also searched

Report this job

⚑ Quick Apply

Create your account and upload your CV to apply for β€” takes less than a minute.

✨ Get a free AI ATS Score Report for your CV the moment you sign up.