About this role
About the Role Canonical is seeking a Site Reliability Engineer to join our globally distributed team and help design, implement, and operate reliable, scalable infrastructure that powers our open source software and platforms. This role supports Ubuntu, cloud initiatives, data science, AI, engineering innovation, and IoT initiatives crafted for enterprise customers worldwide.
What You'll Do
- Design, build, and operate scalable, highly available services and infrastructure
- Monitor system health, respond to incidents, and drive incident response improvements
- Implement automation and strong configuration management (IaC) to reduce toil
- Collaborate with platform and product teams across regions to deliver reliable experiences
- Participate in on-call rotation and post-incident reviews to continuously improve
What We're Looking For
- Experience in site reliability, SRE or related roles; strong Linux/Unix fundamentals
- Proficiency with cloud platforms (AWS/GCP/Azure) and container orchestration (Kubernetes, Docker)
- Automation and scripting skills (e.g., Python, Bash, Terraform, Ansible)
- Observability, monitoring, metrics, logging, and incident management expertise
- Ability to work effectively in a distributed, remote-friendly environment
Nice to Have
- Experience with Canonical technologies (Ubuntu, Juju, MAAS, OpenStack) or open source contributions
- Experience with AI/ML workloads or data-intensive applications
- Security best practices and compliance experience