←  العودة إلى كل الوظائف

Site Reliability Engineer

Thales

قطاع: Technology & IT

📍 الإمارات
💼 دوام كامل
🕒 نُشرت قبل 3 أسابيع

وصف الوظيفة

Role Overview

Site Reliability Engineer (DevSecOps / Sovereign Cloud Engineer) at Thales in Khobar, Saudi Arabia.

Company Overview

Thales is a global digital security company serving business and governments across identity management, data protection, and cybersecurity. More than 30,000 organizations rely on Thales to verify identities, grant access to digital services, analyze information, and encrypt data. For over 50 years, Thales has been a trusted partner of the Kingdom of Saudi Arabia, providing advanced solutions in Defence, Public Security, Civil Aviation, Space, Enterprise, and Cybersecurity across four sites. Thales established SAMI Thales Electronic Systems (STES), a Joint Venture with Saudi Arabian Military Industries (SAMI), to localize defence technologies and build capabilities aligned with Saudi Vision 2030. Thales employs 80,000 employees in 68 countries with a mobility policy supporting career development globally.

Role Purpose

Design, operate, and secure regulated cloud environments across AWS and GCP by combining cloud operations, security engineering, and SRE practices. Drive security incident response, platform reliability, observability, and 24x7 incident management while embedding security-by-design principles, automation-first practices, and continuous improvement into sovereign cloud platforms.

Key Responsibilities

Cloud Operations & Strategy

  • Deploy, manage, and maintain the organization's Sovereign Cloud strategy, ensuring compliance with regulatory and data residency requirements on AWS/GCP public cloud.
  • Manage and operate Kubernetes clusters including upgrades, scaling, and workload optimization.
  • Support and enhance the organization's AWS cloud strategy by embedding security best practices and continuously improving the cloud security posture.

Infrastructure as Code & Automation

  • Implement and manage Infrastructure as Code (Terraform) to provision, modify, and secure cloud resources.
  • Maintain and optimize CI/CD pipelines using Git/GitLab, ensuring secure and automated deployments.
  • Develop automation using Python/Bash/Terraform/Ansible to reduce manual effort, improve operational efficiency, and strengthen platform resilience.

Security & Compliance

  • Investigate, analyze, and remediate cloud security incidents, proactively identifying and mitigating vulnerabilities within AWS environments.
  • Implement and tune security tools, automate policy-driven responses, and advocate DevSecOps practices to ensure secure-by-design cloud operations.
  • Manage secrets and secure access using HashiCorp Vault, including token lifecycle, access policies, and secrets rotation.
  • Conduct proactive system hardening, vulnerability remediation, performance tuning, and capacity planning across cloud environments.

Observability & Reliability

  • Monitor system health and performance using Datadog, define and manage SLIs/SLOs, and drive continuous reliability improvements aligned with SRE principles.
  • Troubleshoot complex infrastructure, networking, container, and performance issues across distributed systems.

Incident Management & Governance

  • Manage 24x7 alerting and incident response through PagerDuty.
  • Lead incident bridges during P1/P2 outages, coordinating cross-functional teams and driving timely resolution with clear root cause analysis.
  • Perform root cause analysis and actively contribute to incident, problem, and change management processes.

Qualifications & Experience

  • 4–5 years of hands-on experience in DevSecOps or SRE roles.
  • Proven hands-on expertise managing Kubernetes clusters.
  • Experience with Terraform (Infrastructure as Code) deployment.
  • Experience with Git/GitLab for CI/CD and version control.
  • Experience working with HashiCorp Vault, Terraform, Datadog, PagerDuty, and Confluence.
  • Experience in incident management, change management, and root cause analysis processes.
  • Strong understanding of SRE principles including SLIs, SLOs, error budgets, and reliability metrics.
  • Hands-on scripting experience in Python and/or Bash to drive automation-first practices.
  • Solid grasp of networking fundamentals (TCP/IP, DNS, load balancing, firewalls, VPNs, private endpoints).
  • Experience or exposure to regulated and compliance-driven environments.
  • Cloud certifications are a plus.

Skills & Competencies

  • Strong understanding of Cloud Security principles including IAM, encryption, network security, container security, and vulnerability management.
  • Linux mandatory.
  • Kubernetes mandatory.
  • Terraform mandatory.
  • Ansible mandatory.
  • Google Cloud mandatory.
  • Python and/or Bash scripting.
  • Maintains a strong customer and business-focused mindset while prioritizing tasks.
  • Cross-functional incident leadership and coordination.

Additional Information

  • Thales provides careers and mobility opportunities, with thousands of employees developing careers globally each year at home and abroad in existing areas of expertise or new fields.

الباحثون عن هذه الوظيفة بحثوا أيضاً عن

الإبلاغ عن هذه الوظيفة

⚡ تقدّم سريع

أنشئ حسابك وارفع سيرتك الذاتية للتقدّم إلى — يستغرق أقل من دقيقة.

✨ احصل على تقرير تقييم مجاني بالذكاء الاصطناعي لسيرتك الذاتية فور التسجيل.