Job description
Role Overview
Expert Site Reliability Engineer at TAWANTECH. This role drives the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.
Role Purpose
To establish and maintain advanced reliability engineering practices across critical technology services, ensuring optimal system performance, incident resilience, and operational efficiency through automation, observability, and technical leadership.
Key Responsibilities
Reliability Engineering & Standards
- Define and implement advanced reliability engineering practices across critical technology services.
- Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
- Identify reliability risks and recommend architectural and engineering improvements.
- Promote automation and engineering approaches that reduce operational toil and improve service reliability.
Automation & Infrastructure
- Design automation to reduce manual operational activities and improve system resilience.
- Drive performance engineering and capacity planning for critical services.
- Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
Incident Management & Analysis
- Lead technical analysis and resolution of complex production incidents.
- Conduct root-cause analysis and drive permanent corrective and preventive actions.
System Design & Resilience
- Design solutions to improve system availability, scalability, capacity, and disaster resilience.
Technical Leadership
- Provide advanced technical guidance and mentorship on SRE practices.
Qualifications & Experience
- Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
- Strong experience in cloud platforms, Kubernetes, and production environments.
- Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).
- Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
- Experience driving reliability improvements and reducing operational toil through automation.
- Experience in Banking, FinTech, or Payment environments is preferred.
Skills & Competencies
- Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
- Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
- Strong analytical, problem-solving, and technical leadership skills.