About this role
Senior Network Production Engineer, AI Supercomputing
United Kingdom, London, London
We are building and operating frontier-scale AI supercomputers used to train the world’s most advanced models. The backend network is a critical part of the machine: a single degraded NIC, switch, link, or path can reduce training performance or disrupt jobs spanning thousands of GPUs.
We are looking for a senior, hands-on Network Production Engineer to own the production health and operations of our multi-rail Ethernet backend network built on MRC and RDMA technologies. You will bridge network engineering, distributed systems, hardware health, and production operations to ensure the network delivers predictable performance at extreme scale.
Your mission is simple: make the network invisible to researchers by maximizing large-job success, minimizing performance degradation, and recovering quickly when failures occur.
Responsibilities
-
Own Production Network Health
-
Own the availability, performance, and operational readiness of the MRC Ethernet backend network.
-
Define and operate service-level indicators for packet loss, congestion, link health, path diversity, bandwidth, tail latency, collective performance, and job impact.
-
Detect slow or degraded components before they cause training failures or reduce model FLOPs utilization.
-
Build health models that correlate switch, NIC, host, topology, and application telemetry.
-
Operate the Network at Frontier Scale
-
Participate in on-call rotation and lead the response to high-severity network incidents.
-
Diagnose failures spanning GPUs, NICs, switches, cables, firmware, drivers, network operating systems, MRC, NCCL, Kubernetes, and training workloads.
-
Develop automated mitigation mechanisms, including path avoidance, node quarantine, workload relocation, and safe component remediation.
-
Create operational procedures for maintenance, upgrades, rollback, capacity expansion, and topology changes.
-
Improve Large-Job Reliability and Performance
-
Work directly with training teams to investigate collective-performance degradation, stalls, timeouts, job restarts, and unexplained MFU loss.
-
Translate low-level network signals into clear job-level impact.
-
Establish network qualification gates for admitting nodes and racks into large training pools.
-
Run failure-injection and resilience testing to validate behavior under link, NIC, switch, and path failures.
-
Improve placement and routing policies for jobs spanning racks, rows, and network planes.
-
Automate Fleet Operations
-
Build production-quality software and automation for network validation, monitoring, diagnosis, remediation, and fleet-wide change management.
-
Replace manual investigation with deterministic workflows and actionable alerts.
-
Automate firmware, driver, configuration, and network operating-system compliance checks.
-
Maintain authoritative network inventory, topology, and configuration state.
-
Reduce mean time to detect, isolate, mitigate, and permanently resolve failures.
-
Drive Cross-Company Execution
-
Lead technical investigations involving internal infrastructure teams, cloud and datacenter operators, Azure Networking, NVIDIA, switch and NIC teams, and netwo