Talent Apply
Log in
All jobs
D

Senior Software Engineer, AI Runtime

databricks
Mountain View, California; San Francisco, California
On-siteUSD 160,000 - 225,000 / year

About this role

About the Role Databricks is building and operating AIR (AI Runtime), a managed platform for large-scale GPU training and fine-tuning. As a Senior Software Engineer for AIR, you will help shape the architecture, scale, and developer experience of the GPU training stack, spanning scheduling, distributed training performance, fault tolerance, and multi-node operations across thousands of GPUs.

What You'll Do

  • Drive the architecture and evolution of AIR's managed GPU training platform for scalable, high-throughput, and resilient training across fleets of accelerators.
  • Solve hardest problems in large-scale training: multi-node orchestration, distributed parallelism (data, tensor, pipeline, sequence), GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint/restore for long-running jobs.
  • Improve GPU efficiency and training performance by raising utilization and lowering cost per training run across model architectures and hardware generations.
  • Build resilience and observability foundations to keep multi-node jobs healthy, detecting and recovering from failures with minimal customer disruption.
  • Partner with product, research, and platform teams to shape APIs, CLI, and the developer experience for launching, monitoring, and debugging production training jobs.
  • Lead end-to-end engineering efforts from design through production rollout with high standards for performance, correctness, and reliability.
  • Make direct contributions to core AIR systems and support for latest accelerators and new regions as the fleet grows.
  • Champion engineering excellence, mentor other engineers, and contribute to Databricks' technical direction in AI training infrastructure.

What We're Looking For

  • 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, HPC, or ML systems.
  • Experience with distributed training frameworks (PyTorch, FSDP, DeepSpeed, Megatron) and parallelism strategies (data, tensor, pipeline, sequence).
  • Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.
  • Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (NVLink and InfiniBand or RoCE), collective communication, and bottlenecks that govern throughput and utilization.
  • Experience building and operating managed, multi-tenant cloud platform products with clear SLAs and SLOs.
  • Strong foundation in algorithms, data structures, and system design for performance-sensitive, large-scale distributed systems.
  • Proven ability to deliver technically complex, high-impact initiatives with clear customer or business value.
  • Strong communication skills and ability to collaborate across product, research, and infrastructure teams in a fast-moving environment.
  • Customer-focused mindset with ability to align implementation details with product goals, and a passion for mentoring engineers and fostering technical excellence.
  • BS in Computer Science or related field (MS or PhD preferred).

Compensation & Benefits

  • Local Pay Range: USD 160,000—USD 225,000 per year. The total compensation package may include annual performance bonus, equity, and listed benefits. For location-specific ranges, see the page linked in the posting.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →