About this role
Job title: Staff Software Engineer, AI Runtime
About the Role As a Staff Software Engineer for AI Runtime (AIR), you will help design, implement, and scale the managed GPU training stack that powers large-scale model training and fine-tuning. You will influence architecture, performance, reliability, and developer experience for multi-node, multi-accelerator workloads across Databricks' Mosaic AI mission.
What You'll Do
- Drive the architecture and evolution of AIR's managed GPU training platform to deliver scalable, high-throughput, and resilient training across fleets of thousands of accelerators.
- Solve hard problems in large-scale training: multi-node orchestration, distributed parallelism strategies (data, tensor, pipeline, sequence), GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint/restore for long-running jobs.
- Improve GPU efficiency and training performance (utilization, end-to-end throughput) while lowering cost per training run across model architectures and hardware generations.
- Build resilience and observability foundations to keep multi-node jobs healthy, with fast failure detection and minimal user disruption.
- Partner with product, research, and platform teams to shape APIs, CLI, and the developer experience for launching, monitoring, and debugging production training jobs.
- Lead end-to-end engineering efforts from design to production rollout, maintaining high standards for performance, correctness, and reliability.
- Contribute to core AIR systems and bring up support for new accelerators and regions as the fleet grows.
- Mentor engineers, review designs, and help shape Databricks' long-term technical direction in AI training infrastructure.
What We're Looking For
- 10+ years building and operating large-scale distributed systems, with deep experience in GPU training infrastructure, HPC, or ML systems.
- Hands-on experience with distributed training frameworks (PyTorch, FSDP, DeepSpeed, Megatron) and parallelism strategies (data, tensor, pipeline, sequence).
- Strong understanding of training resilience (checkpointing, failure detection, automatic recovery) for long-running multi-node jobs.
- Solid knowledge of GPU performance fundamentals (accelerator architecture, NVLink/InfiniBand or RoCE, collective communication, throughput bottlenecks).
- Experience delivering managed, multi-tenant cloud platform products with clear SLAs/SLOs.
- Strong algorithms, data structures, and system design for performance-sensitive distributed systems.
- Proven ability to deliver technically complex, high-impact initiatives with clear customer/business value.
- Excellent communication and cross-team collaboration skills; strategic, product-oriented mindset; passion for mentoring.
- BS in Computer Science or related field; MS or PhD preferred.
Compensation & Benefits
- Local Pay Range: $190,000–$265,000 USD per year.
- This role may include annual performance bonus, equity, and the benefits listed above. Compensation varies by location and other factors.