وصف الوظيفة
Role Overview
Senior Platform Engineer - Cloud & Kubernetes at Weekday AI, located in Abu Dhabi, United Arab Emirates. This is a full-time position for one of the company's clients.
Role Purpose
Design, implement, secure, and operate scalable cloud-native platforms using Microsoft Azure and Kubernetes. Build highly available, resilient, secure, and well-governed production environments while driving automation, observability, infrastructure-as-code, and platform engineering best practices.
Key Responsibilities
Infrastructure Design & Deployment
- Design, deploy, and manage Microsoft Azure infrastructure and AKS/Kubernetes platforms.
- Implement and maintain Disaster Recovery, backup, restore, and business continuity solutions.
- Plan and execute platform migrations, infrastructure upgrades, and cloud transformation initiatives.
Automation & Infrastructure-as-Code
- Build infrastructure automation using Terraform, Helm, Azure DevOps, and CI/CD pipelines.
- Define Infrastructure-as-Code and CI/CD governance, including reusable Terraform modules, state management, code reviews, approval gates, environment promotion, and artifact management.
- Establish reusable Infrastructure-as-Code patterns and platform engineering standards.
Kubernetes & Platform Configuration
- Manage Kubernetes networking, ingress, storage, namespaces, resource limits, security, and platform configurations.
- Manage TLS certificates, secrets, access controls, RBAC, and platform security mechanisms.
Observability & Monitoring
- Implement observability solutions using Prometheus, Grafana, Loki, and related monitoring technologies.
- Define and monitor platform availability, SLIs, SLOs, capacity, performance, and operational health.
Governance & Security
- Establish cloud and Kubernetes governance covering resource organisation, naming, tagging, security policies, and operational standards.
- Drive security and compliance readiness through vulnerability remediation, access reviews, security baselines, audit controls, and policy enforcement.
- Strong understanding of Kubernetes security, governance, RBAC, policy enforcement, and Azure Policy.
Operations & Incident Management
- Lead incident and problem management, including root-cause analysis, corrective actions, and prevention of recurring issues.
- Own DR testing, RTO/RPO validation, backup and recovery standards, and periodic recovery exercises.
- Perform capacity planning, performance optimisation, and cloud cost optimisation / FinOps activities.
Architecture & Technical Review
- Participate in architecture and technical design reviews for platforms, applications, integrations, and infrastructure changes.
- Strong experience participating in architecture and technical design reviews.
Technology Evaluation & Innovation
- Evaluate emerging platform technologies, conduct POCs, and establish approved patterns for production adoption.
Documentation & Knowledge Management
- Maintain technical documentation including HLDs, LLDs, architecture diagrams, SOPs, runbooks, troubleshooting guides, and DR procedures.
- Provide technical leadership, mentoring, and knowledge sharing across platform engineering teams.
Qualifications & Experience
- 5+ years of experience in platform engineering, cloud infrastructure, DevOps, SRE, or related roles.
- Strong hands-on expertise in Microsoft Azure and Kubernetes, particularly AKS.
- Strong experience with Docker, Helm, Terraform, and Azure DevOps / CI/CD.
- Good understanding of Azure networking, including VNets, NSGs, Private Endpoints, and Azure Firewall.
- Experience implementing observability using Prometheus, Grafana, Loki, or similar platforms.
- Proven experience with Disaster Recovery, backup/restore, RTO/RPO planning, and recovery testing.
- Experience working with databases such as PostgreSQL, MongoDB, MySQL, or Azure SQL.
- Experience working in security-, compliance-, audit-, or governance-controlled environments.
- Experience in banking or other regulated industries would be an advantage.
Skills & Competencies
- Strong knowledge of Linux, Bash scripting, Git, and GitOps practices.
- Strong knowledge of incident management, problem management, RCA, capacity planning, performance optimisation, and cloud cost management.
- Ability to create and maintain HLDs, LLDs, architecture diagrams, operational runbooks, SOPs, and technical documentation.
- Excellent troubleshooting, analytical, and problem-solving capabilities.
- Strong technical leadership, mentoring, communication, and cross-functional collaboration skills.
- Strong understanding of high availability, production operations, platform modernisation, and cloud transformation initiatives.
Additional Information
Salary range: INR 26–45 LPA (Rs 26,000,000 - Rs 45,000,000).