Talent Apply
Log in
All jobs
UE

Site Reliability Specialist

Ubisoft Entertainment
Montreal, Canada (Office-based)
On-site

About this role

Job title: Site Reliability Specialist

Ubisoft Montréal is seeking a Site Reliability Specialist to join the IT Games and Studios team to improve availability, reliability, and performance of critical platforms and services that support game development across Ubisoft. You will collaborate with developers, cloud specialists, and infrastructure teams to build resilient solutions and support operational excellence across production environments.

What You'll Do

  • Collaborate with service teams to define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Design and implement automation solutions that improve operational efficiency and service reliability
  • Document technical solutions and support their integration across internal platforms and services
  • Align technical implementations with established quality standards, engineering practices, and IT guidelines
  • Partner with development, infrastructure, and platform teams to improve operational consistency and system reliability
  • Support observability practices, including monitoring, logging, alerting, and incident management
  • Contribute to root cause analysis and continuous improvement initiatives following service incidents
  • Optimise deployment workflows and operational processes through automation and infrastructure improvements
  • Support the maintenance and evolution of cloud-based environments and services
  • Share knowledge and contribute to reliability-focused initiatives across teams

What We're Looking For

  • Experience with infrastructure engineering, automation, and DevOps practices
  • Knowledge of GitLab and GitLab CI/CD for deployment and automation workflows
  • Proficiency with scripting and programming languages such as Python, Bash, and Go
  • Experience using Terraform, Infrastructure as Code (IaC) practices, and Kubernetes (K8s) in public cloud environments such as Amazon Web Services (AWS) or Microsoft Azure
  • Familiarity with configuration management tools such as Ansible or Chef
  • Experience with observability and monitoring platforms such as Prometheus and Grafana
  • Ability to collaborate effectively with technical and non-technical partners while supporting problem-solving and continuous improvement initiatives

Nice to Have

  • Interest in AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar solutions that support development, automation, and operational efficiency

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →