وصف الوظيفة
Role Overview
Principal SRE at Alpheya, a wealth-management technology company headquartered in Abu Dhabi. You will report to the CTO and own SaaS operations end to end, from Kubernetes clusters to quarterly service reviews with bank executives.
Company Overview
Alpheya is a wealth-management technology company that builds tailored investing experiences for banks—from mobile apps to advisor and back-office portals—on a single shared platform covering the full order-to-custody lifecycle. The platform runs as SaaS on Microsoft Azure, with on-premises delivery for banks that require it.
Role Purpose
You are accountable for the end-to-end service that customers buy: SLAs, incident reviews with bank CIOs, audits, disaster recovery, and the cost of running each tenant. Your mandate is to make onboarding the fifth bank a checklist instead of a project. Over the next six months, the company is taking multiple banks live as SaaS customers, and you will establish the operational discipline required to scale.
Key Responsibilities
Service Management
- Own service management for every SaaS customer, including SLA definition and reporting, incident communications, and post-incident reviews.
- Manage security questionnaires and audit cycles.
- Establish the day-to-day support model with each bank's service desk.
Release and Deployment Operations
- Own the release calendar across the tenant estate.
- Establish environment promotion and rollback discipline.
- Manage production change management and coordinate rollouts with each bank's change and freeze windows.
- Decide when and how production changes are deployed once engineering builds the pipelines.
Reliability and Capacity
- Lead the reliability program, including tested disaster recovery and business continuity.
- Own backup verification, patching and vulnerability-management cadence.
- Develop capacity planning and a cost-per-tenant model for the CFO.
Onboarding and Expansion
- Establish the tenant onboarding runbook so new banks go live on a repeatable path.
- Operate the European launch with two new Azure regions, including data residency, DR, and support coverage.
Compliance and Risk
- Manage compliance operations for outsourcing arrangements, working with bank risk teams under frameworks such as ISO 27001, SOC 2, and European and Gulf outsourcing regulation.
Qualifications & Experience
- 12+ years in production operations, several of them leading the function for multi-tenant SaaS.
- Proven ownership of customer-facing service management: led severity-one incidents, presented post-incident reviews to customer executives, and managed customer audits.
- Built or scaled an SRE or platform-operations team and designed on-call coverage across regions.
- Working knowledge of Azure: regions and availability zones, networking and private connectivity, identity, and quota planning.
- Experience with certification and audit regimes: ISO 27001, SOC 2, or bank outsourcing regulation in the EU or the Gulf.
Skills & Competencies
- Deep Kubernetes and cloud expertise to challenge engineers on failure modes, disaster recovery design, and capacity claims.
- Ability to negotiate infrastructure decisions with Microsoft and bank security teams.
Additional Information
First Six Months
- Deliver four bank go-lives across two geographies, each with agreed SLAs, escalation paths, and incident procedures in place from day one.
- Run and document a disaster-recovery exercise for at least one production environment.
- Establish an on-call rotation covering both regions without heroics.
- Develop a tenant cost model and capacity plan for the whole estate.
Nice to Have
- Wealth management, brokerage, or capital-markets domain exposure.
- Experience operating streaming or workflow infrastructure such as Temporal or Kafka.
- Experience with on-premises software delivery to enterprise customers.