وصف الوظيفة
Role Overview
10x Software Engineer at Lucidya, an AI-powered Customer Experience Management platform in the MENA region. This is a full-stack backend engineering role focused on distributed systems, platform reliability, and scale.
Company Overview
Lucidya is a leading AI-powered Customer Experience Management (CXM) platform in the MENA region, enabling enterprises to understand, engage, and serve customers across digital channels at scale. The company is moving toward IPO-scale and rebuilding core platform components to achieve extreme reliability, high-scale distributed processing of billions of data points, and AI-native architecture incorporating LLM and real-time intelligence.
Role Purpose
You will design, build, and operate a high-throughput, event-driven microservice platform serving billions of data points. You will drive platform decoupling, eliminate root causes of production failures, and build systems with extreme reliability, observability, and performance. The role demands extreme ownership: you fix broken systems regardless of team boundaries, and you think deeper and move faster than the average engineer.
Key Responsibilities
Microservices & Distributed Architecture
- Design and operate high-throughput, event-driven pipelines across a 100+ microservice ecosystem handling billions of data points.
- Build and scale distributed messaging systems with RabbitMQ, including backpressure management, consumer scaling, and queue health monitoring.
- Develop and maintain API gateway layers with advanced routing capabilities: multi-upstream support, traffic splitting, and environment isolation.
- Architect SSO and identity federation for enterprise clients, supporting multi-IdP routing with zero coupling to core services.
- Define clean service boundaries across ingestion, processing, and delivery pipelines spanning Ruby and Python.
Performance & Systems
- Diagnose and resolve complex production issues such as deadlocks, queue exhaustion, and connection pool saturation, eliminating root causes.
- Optimize PostgreSQL for heavy write workloads, contention management, schema design, triggers, and connection scaling.
- Design and tune Elasticsearch for search, indexing, and real-time Arabic relevance at scale.
- Make informed trade-offs between multi-process and async architectures based on workload characteristics.
Observability & Reliability
- Build and maintain observability across a large-scale system using Grafana, Loki, distributed tracing, and SLOs.
- Own production incidents end-to-end, tracing failures across queues, search systems, and external integrations.
- Lead root cause analysis and implement preventative measures across multi-service pipelines.
- Build internal tooling that improves engineering velocity, automation, deployment gating, and review enforcement.
- Turn architectural principles into enforceable standards and guardrails, not just documentation.
Platform Evolution
- Drive platform decoupling and service isolation across the system.
- Contribute to Kubernetes migration and infrastructure modernization.
- Standardize and improve CI/CD pipelines across services.
Qualifications & Experience
- Strong foundation in distributed systems; you understand failure modes before writing the first line.
- Hands-on production experience with event-driven architecture and message queues.
- Track record debugging and preventing complex production issues, not just fixing them.
- Experience with Rails or Python backends at meaningful scale.
- Demonstrated ability to improve systems you were not asked to touch.
Skills & Competencies
- Deep comfort with concurrency, backpressure, and fault tolerance.
- Proficiency with Ruby on Rails and/or Python.
- Experience with PostgreSQL, Elasticsearch, Redis, and RabbitMQ.
- Working knowledge of Kubernetes, AWS or GCP, and APISIX.
- Strong observability mindset: Grafana, Loki, and distributed tracing.
- Ability to read and understand existing codebases rapidly.
- Strong opinions about architecture backed by data.
- Systems thinking: latency, throughput, failure modes, and cost at scale.
- Treat documentation, tests, and observability as non-negotiable defaults.
- Ship fast without breaking things; speed and quality are not a trade-off.
- Sense of urgency that does not require external pressure.
Additional Information
- You will have real scale with billions of events, not toy systems.
- Direct impact at CTO and executive level.
- Pre-IPO company with clear growth trajectory; your work has real impact on clients.