Talent Apply
Log in
All jobs
M

AI Reliability Engineer, HPC

Microsoft
London, United Kingdom Posted Oct 8, 2026
On-site

About this role

AI Reliability Engineer, HPC

United Kingdom, London, London

Job description

Company and benefits

Job number

200060491

Date posted

Oct 08, 2026

Work site 4 days / week in-office

Travel Less than 25%

Profession Software Engineering

Discipline Software Engineering

Role type Individual Contributor

Employment type Full-Time

Overview

As Microsoft continues to push the boundaries of AI, we are on the lookout for passionate individuals to work with us on the most interesting and challenging AI questions of our time. Our vision is bold and broad — to build systems that have true artificial intelligence across agents, applications, services, and infrastructure. It’s also inclusive: we aim to make AI accessible to all — consumers, businesses, developers — so that everyone can realize its benefits.

We’re looking for an experienced AI Reliability Engineer to join our High Performance Computing (HPC) infrastructure team. In this role, you’ll blend software engineering and systems engineering to keep our large-scale distributed AI infrastructure reliable and efficient. You’ll ensure that AI systems stay efficient and reliable with very high uptimes.

Microsoft AI Our mission is to build AI that amplifies human potential and empowers people around the world. We strive to deliver breakthroughs that advance science, education, productivity, and global well-being. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models! MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.

Responsibilities

  • Reliability & Availability: Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference.
  • Observability: Design and maintain monitoring, alerting, and logging systems to provide real-time visibility into all aspects of HPC systems including GPU, clusters, storage and networking.
  • Automation & Tooling: Build automation for deployments, incident response, scaling, and failover in CPU+GPU environments.
  • Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements.
  • Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments.
  • Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows.
  • Embody our Culture and Values.

Qualifications

Required Qualifications:

  • Bachelor’s Degree in Computer Science, or related technical discipline AND 4+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering OR equivalent experience

Preferred Qualifications:

  • Master’s Degree in Computer Science, or related technical discipline AND 2+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering
  • OR equivalent experience
  • Experience with Kubernetes, Docker, container orchestration, and CI/CD pipelines for ML training or inference workloads.
  • Experience with public cloud platforms such as Azure, AWS, or GCP, including infrastructure-as

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →