About this role
About the Role Databricks is building and operating AIR (AI Runtime), a managed platform for large-scale GPU training and fine-tuning. As a Senior Software Engineer for AIR, you will help shape the architecture, scale, and developer experience of the GPU training stack, spanning scheduling, distributed training performance, fault tolerance, and multi-node operations across thousands of GPUs.
What You'll Do
- Drive the architecture and evolution of AIR's managed GPU training platform for scalable, high-throughput, and resilient training across fleets of accelerators.
- Solve hardest problems in large-scale training: multi-node orchestration, distributed parallelism (data, tensor, pipeline, sequence), GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint/restore for long-running jobs.
- Improve GPU efficiency and training performance by raising utilization and lowering cost per training run across model architectures and hardware generations.
- Build resilience and observability foundations to keep multi-node jobs healthy, detecting and recovering from failures with minimal customer disruption.
- Partner with product, research, and platform teams to shape APIs, CLI, and the developer experience for launching, monitoring, and debugging production training jobs.
- Lead end-to-end engineering efforts from design through production rollout with high standards for performance, correctness, and reliability.
- Make direct contributions to core AIR systems and support for latest accelerators and new regions as the fleet grows.
- Champion engineering excellence, mentor other engineers, and contribute to Databricks' technical direction in AI training infrastructure.
What We're Looking For
- 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, HPC, or ML systems.
- Experience with distributed training frameworks (PyTorch, FSDP, DeepSpeed, Megatron) and parallelism strategies (data, tensor, pipeline, sequence).
- Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.
- Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (NVLink and InfiniBand or RoCE), collective communication, and bottlenecks that govern throughput and utilization.
- Experience building and operating managed, multi-tenant cloud platform products with clear SLAs and SLOs.
- Strong foundation in algorithms, data structures, and system design for performance-sensitive, large-scale distributed systems.
- Proven ability to deliver technically complex, high-impact initiatives with clear customer or business value.
- Strong communication skills and ability to collaborate across product, research, and infrastructure teams in a fast-moving environment.
- Customer-focused mindset with ability to align implementation details with product goals, and a passion for mentoring engineers and fostering technical excellence.
- BS in Computer Science or related field (MS or PhD preferred).
Compensation & Benefits
- Local Pay Range: USD 160,000—USD 225,000 per year. The total compensation package may include annual performance bonus, equity, and listed benefits. For location-specific ranges, see the page linked in the posting.