وصف الوظيفة
Role Overview
Site Reliability Engineer at Lucidya. Lucidya is an AI-native platform for customer experience (CX) intelligence that manages entire customer lifecycles autonomously, from initial engagement through retention and growth.
Company Overview
Lucidya is an AI-native platform for customer experience (CX) intelligence that processes massive volumes of real-time customer data. The platform uses proprietary NLU technology built in-house and trained on millions of multilingual conversations, enabling marketing, support, CX, and research teams to deliver personalized experiences that drive measurable improvements in customer satisfaction, retention, and lifetime value. As the company scales globally, the reliability, performance, and resilience of infrastructure become mission-critical.
Role Purpose
As a Site Reliability Engineer, you will own the reliability of Lucidya's cloud infrastructure and ensure it scales seamlessly as the organization grows. You will anticipate infrastructure challenges, design systems that prevent failures, and build automation that removes issues entirely. Any downtime, latency, or instability directly impacts customers' ability to make decisions and serve their own users; this role ensures that does not happen.
Key Responsibilities
Infrastructure Design & Reliability
- Design and maintain infrastructure that is highly available, fault-tolerant, and scalable.
- Proactively identify and eliminate single points of failure before they become incidents.
- Ensure production systems remain stable, even under increasing scale and load.
Cloud Environment Management
- Manage and continuously improve workloads across AWS, GCP, or Azure.
- Use Infrastructure as Code (Terraform) to standardize and scale infrastructure.
- Optimize resource usage to balance performance and cost.
Kubernetes Operations
- Operate and scale Kubernetes clusters (EKS, GKE, etc.) with confidence.
- Troubleshoot issues quickly and ensure smooth deployments and upgrades.
- Ensure containerized workloads perform reliably at scale.
Observability & Incident Management
- Implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK.
- Define alerting that is meaningful, not noisy.
- Respond to incidents, lead root cause analysis, and ensure the organization learns from every failure.
Automation & Efficiency
- Write scripts and build tooling to eliminate repetitive operational work.
- Continuously improve infrastructure efficiency through automation.
- Promote a culture where manual work is a temporary state, not the norm.
Collaboration & Best Practices
- Work closely with DevOps and engineering teams to solve performance bottlenecks.
- Contribute to CI/CD improvements and deployment reliability.
- Help shape reliability best practices across the organization.
Qualifications & Experience
- Approximately 3 years of hands-on experience working in SRE, DevOps, or infrastructure engineering, with exposure to systems at scale.
- Hands-on experience working with Kubernetes in production and troubleshooting when issues arise.
- Comfortable working in cloud environments like AWS, GCP, or Azure with understanding of how distributed systems behave.
Skills & Competencies
Infrastructure & Cloud Tools
- Use Terraform or similar Infrastructure as Code tools to manage infrastructure.
- Work confidently with Docker and Kubernetes.
- Understand networking, load balancing, and high-availability design.
- Understand CI/CD pipelines (Jenkins, GitHub Actions, Bitbucket, etc.).
Scripting & Automation
- Write scripts in Python, Bash, or similar languages to automate workflows.
Monitoring & Observability
- Implement tools like Prometheus, Grafana, Datadog, or ELK.
- Know the difference between useful alerts and noise.
- Focus on signals that actually drive action.
Personal Qualities
- Take ownership—do not wait to be told something is broken.
- Remain calm under pressure and methodical during incidents.
- Simplify complexity instead of adding to it.
- Communicate clearly, even when explaining deeply technical issues.
- Care about building systems that make other engineers more effective.
Additional Information
Nice to Have
- Experience with RabbitMQ or Redis in production.
- Familiarity with Ansible or AWX.
- Exposure to multi-cloud or hybrid environments.
- Cloud certifications (AWS, GCP) or Linux certifications.
- Background from ITI (Information Technology Institute).
First 90 Days
By day 30, you will have built a strong understanding of Lucidya's infrastructure, systems, and workflows; contribute to day-to-day operations with support from the team; and start identifying areas for improvement in automation and reliability.
By day 90, you will independently manage infrastructure tasks and troubleshoot issues, actively contribute to reliability and scalability improvements, and take ownership of parts of the infrastructure while improving them.
Hiring Process
- Screening Interview – Talent Acquisition.
- Technical Interview – SRE Lead.
- Technical Task.
- Final Interview – SRE Lead & Cloud DevOps Director.