About this role
Job title: Site Reliability Engineering Manager
About the Role A Site Reliability Engineering Manager is sought to lead Canonical's APAC SRE efforts from a remote location. The role focuses on reliability, availability and performance of Canonical's platforms, including Ubuntu-based services used by global customers.
What You'll Do
- Lead and grow a high-performing SRE team in APAC, setting strategic direction and priorities.
- Define and drive SRE best practices, incident response, blameless postmortems and post-incident reviews.
- Collaborate with product, engineering and customer-facing teams to improve reliability and scalability of platforms and services.
- Build and automate infrastructure, monitoring, and lifecycle management to reduce toil and improve observability.
- Manage on-call rotation, SLAs/SLOs, and capacity planning to meet business needs.
What We're Looking For
- Proven experience in site reliability engineering or production engineering with leadership experience.
- Strong background in Linux systems, cloud platforms, Kubernetes, monitoring (Prometheus, Grafana, etc.), and automation.
- Demonstrated incident management and postmortem culture.
- Excellent communication skills and ability to work across time zones in APAC.
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).
Nice to Have
- Experience with open-source ecosystems and Canonical's stack (Ubuntu, cloud-native tooling).
- Prior experience managing distributed teams across APAC.
- Familiarity with security, compliance and governance in cloud environments.
Compensation & Benefits
- Competitive salary and benefits package; details discussed during interview.
- Flexible work arrangements and remote-first culture.