وصف الوظيفة
Role Overview
Assistant Manager – Cloud Operations at Qiddiya Investment Company. The role supports day-to-day operation, reliability, governance, security, and optimization of Azure and Google Cloud Platform (GCP) environments.
Role Purpose
Support the stable, secure, resilient, and cost-effective operation of cloud services aligned with business requirements. Work across infrastructure, security, application, service desk, vendors, managed service providers, and business stakeholders to ensure cloud environments meet operational excellence and Site Reliability Engineering (SRE) standards.
Key Responsibilities
Cloud Operations & Monitoring
- Support daily cloud operations across Microsoft Azure and Google Cloud Platform (GCP).
- Monitor cloud infrastructure health, availability, performance, capacity, and incidents.
- Track operational KPIs such as availability, incident resolution, SLA compliance, policy violations, cost optimization, reliability metrics, and documentation coverage.
Incident & Service Management
- Manage incident response, service requests, escalations, and root cause analysis for cloud services.
- Participate in change management, release coordination, and operational readiness reviews.
Site Reliability Engineering
- Support SRE practices, including reliability monitoring, availability improvement, incident response, service health reviews, and operational automation.
Governance, Security & Compliance
- Support implementation and maintenance of cloud governance, policies, standards, tagging, and compliance controls.
- Support cloud security operations, including IAM reviews, network security, patching, vulnerability remediation, encryption, and access governance.
- Support audits, compliance reviews, and risk remediation related to Azure and GCP cloud infrastructure.
Cost Optimization
- Review and optimize cloud cost, usage, reserved capacity, rightsizing, and resource cleanup.
Resilience & Business Continuity
- Support native backup, disaster recovery, high availability, and business continuity activities.
Collaboration & Coordination
- Coordinate with vendors, managed service providers, and internal teams to resolve operational issues.
Documentation & Automation
- Maintain cloud documentation, operational runbooks, architecture diagrams, and knowledge base articles.
- Assist with automation and Infrastructure as Code adoption using tools such as Terraform.
Qualifications & Experience
- Minimum 5 years of hands-on experience in cloud operations, cloud engineering, infrastructure operations, SRE, or related technical roles.
- Hands-on knowledge of Microsoft Azure and/or Google Cloud Platform (GCP).
- Understanding of cloud networking, identity, compute, storage, monitoring, backup, security, and high availability.
- Experience with incident, problem, change, and service request management.
- Familiarity with SRE concepts, including availability, reliability, monitoring, alerting, incident response, root cause analysis, and continuous improvement.
- Familiarity with ITIL processes and IT service management tools.
- Good understanding of cloud security, IAM, RBAC, network security groups/firewalls, encryption, vulnerability management, and compliance controls.
Skills & Competencies
- Strong troubleshooting, communication, documentation, and stakeholder coordination skills.
- Ability to work with vendors, managed service providers, and cross-functional teams.
- Experience with automation or scripting using tools such as Terraform, PowerShell, or Python is preferred.