Talent Apply
Log in
All jobs
D

Staff Software Engineer, AI Runtime

databricks
Mountain View, California; San Francisco, California
On-siteUSD 190,000 - 265,000 / year

About this role

Job title: Staff Software Engineer, AI Runtime

About the Role As a Staff Software Engineer for AI Runtime (AIR), you will help design, implement, and scale the managed GPU training stack that powers large-scale model training and fine-tuning. You will influence architecture, performance, reliability, and developer experience for multi-node, multi-accelerator workloads across Databricks' Mosaic AI mission.

What You'll Do

  • Drive the architecture and evolution of AIR's managed GPU training platform to deliver scalable, high-throughput, and resilient training across fleets of thousands of accelerators.
  • Solve hard problems in large-scale training: multi-node orchestration, distributed parallelism strategies (data, tensor, pipeline, sequence), GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint/restore for long-running jobs.
  • Improve GPU efficiency and training performance (utilization, end-to-end throughput) while lowering cost per training run across model architectures and hardware generations.
  • Build resilience and observability foundations to keep multi-node jobs healthy, with fast failure detection and minimal user disruption.
  • Partner with product, research, and platform teams to shape APIs, CLI, and the developer experience for launching, monitoring, and debugging production training jobs.
  • Lead end-to-end engineering efforts from design to production rollout, maintaining high standards for performance, correctness, and reliability.
  • Contribute to core AIR systems and bring up support for new accelerators and regions as the fleet grows.
  • Mentor engineers, review designs, and help shape Databricks' long-term technical direction in AI training infrastructure.

What We're Looking For

  • 10+ years building and operating large-scale distributed systems, with deep experience in GPU training infrastructure, HPC, or ML systems.
  • Hands-on experience with distributed training frameworks (PyTorch, FSDP, DeepSpeed, Megatron) and parallelism strategies (data, tensor, pipeline, sequence).
  • Strong understanding of training resilience (checkpointing, failure detection, automatic recovery) for long-running multi-node jobs.
  • Solid knowledge of GPU performance fundamentals (accelerator architecture, NVLink/InfiniBand or RoCE, collective communication, throughput bottlenecks).
  • Experience delivering managed, multi-tenant cloud platform products with clear SLAs/SLOs.
  • Strong algorithms, data structures, and system design for performance-sensitive distributed systems.
  • Proven ability to deliver technically complex, high-impact initiatives with clear customer/business value.
  • Excellent communication and cross-team collaboration skills; strategic, product-oriented mindset; passion for mentoring.
  • BS in Computer Science or related field; MS or PhD preferred.

Compensation & Benefits

  • Local Pay Range: $190,000–$265,000 USD per year.
  • This role may include annual performance bonus, equity, and the benefits listed above. Compensation varies by location and other factors.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →