←  Back to all vacancies

Expert Site Reliability Engineer

TAWANTECH

Technology & IT

πŸ“ Saudi Arabia
πŸ’Ό Full-time
πŸ•’ Posted 3 weeks ago

Job description

Role Overview

Expert Site Reliability Engineer at TAWANTECH. This role drives the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.

Role Purpose

To establish and maintain advanced reliability engineering practices across critical technology services, ensuring optimal system performance, incident resilience, and operational efficiency through automation, observability, and technical leadership.

Key Responsibilities

Reliability Engineering & Standards

  • Define and implement advanced reliability engineering practices across critical technology services.
  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
  • Identify reliability risks and recommend architectural and engineering improvements.
  • Promote automation and engineering approaches that reduce operational toil and improve service reliability.

Automation & Infrastructure

  • Design automation to reduce manual operational activities and improve system resilience.
  • Drive performance engineering and capacity planning for critical services.
  • Develop and enhance monitoring, observability, alerting, and incident detection capabilities.

Incident Management & Analysis

  • Lead technical analysis and resolution of complex production incidents.
  • Conduct root-cause analysis and drive permanent corrective and preventive actions.

System Design & Resilience

  • Design solutions to improve system availability, scalability, capacity, and disaster resilience.

Technical Leadership

  • Provide advanced technical guidance and mentorship on SRE practices.

Qualifications & Experience

  • Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
  • Strong experience in cloud platforms, Kubernetes, and production environments.
  • Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).
  • Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
  • Experience driving reliability improvements and reducing operational toil through automation.
  • Experience in Banking, FinTech, or Payment environments is preferred.

Skills & Competencies

  • Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
  • Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
  • Strong analytical, problem-solving, and technical leadership skills.

People looking at this role also searched

Report this job

⚑ Quick Apply

Create your account and upload your CV to apply for β€” takes less than a minute.

✨ Get a free AI ATS Score Report for your CV the moment you sign up.