About this role
About the Role
Databricks is seeking an exceptional Senior Staff Technical Program Manager (TPM) for Reliability to lead the strategy, execution, and continuous improvement of our most critical Reliability initiatives across infrastructure and product engineering teams. As Databricks scales for thousands of customers and data-intensive workloads, this role drives cross-company programs that enhance reliability, performance, and operational excellence of our multi-cloud infrastructure. You will partner with senior engineering leaders to define reliability strategy, set long-term goals, and execute multi-quarter programs to build the most reliable cloud platform for mission-critical workloads.
What You'll Do
- Lead Reliability Strategy + Multi-Quarter Roadmaps: define long-term reliability roadmaps, influence technical direction, and align priorities across Platform Engineering, Compute Fleet Management, SRE, Security, and Cloud Partnerships.
- Drive Execution of Critical Reliability Programs: own end-to-end program planning, risk management, dependency mapping, trade-off decisions, status reporting, and delivery; identify gaps and drive improvements.
- Partner Deeply with Engineering & Influence Technical Direction: bring infrastructure, distributed systems, or SRE expertise to help teams make sound design decisions and align on priorities.
- Facilitate Cross-Functional Alignment: ensure programs are technically grounded and execution-ready with strong cross-team collaboration.
- Improve Reliability through Systems Thinking: diagnose bottlenecks and drive improvements in scalability, fault tolerance, automation, and operational tooling.
- Elevate Reliability Culture Across the Organization: promote best practices (error budgets, incident reviews, design-for-resilience, operational readiness) and drive governance, metrics, and scalable processes.
- Evangelize Reliability & Reduce Operational Load: advocate for reliability expectations and engineer-empowering processes that improve incident preparedness.
What We're Looking For
- 10+ years of experience managing and delivering large-scale technical programs in cloud infrastructure, distributed systems, SRE, or platform engineering.
- Experience building infrastructure across two or more hyperscale cloud providers (AWS, Azure, GCP) with knowledge of cloud primitives, multi-AZ/region patterns, and control/data plane concepts.
- Demonstrated success leading Reliability Programs at scale (availability, failover, operational excellence, incident reduction, dependency hardening).
- Strong understanding of infrastructure, distributed systems, or SRE practices; prior engineering or SRE experience is highly preferred.
- Proven ability to partner with senior engineering leadership to define strategy and drive multi-team initiatives.
- Ability to translate ambiguous goals into actionable program plans with milestones, KPIs, and success metrics.
- Experience managing complex cross-organizational dependencies, technical risks, and multi-quarter timelines.
- Experience delivering programs across multiple clouds and/or large-scale cloud-native services.
- Experience building and scaling engineering processes, operational frameworks, and stakeholder alignment mechanisms.
Nice to Have
- Background in distributed systems engineering, SRE, platform infrastructure, or cloud services.
- Experience with large-scale compute fleets, container orchestration, autoscaling, or control-plane architecture.
- Familiarity with reliability methodologies such as SLOs, error budgets, chaos engineering, failure mode analysis, and incident management frameworks.
- Expertise using Jira or equivalent tools for program tracking and execution.
- Bachelor’s degree in Computer Science, Engineering, or related technical field; advanced degree preferred.
Compensation & Benefits
Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role are listed in the posting as part of our transparency efforts. Actual compensation packages are based on several factors including but not limited to experience, skills, certifications, and location. Details on the exact range are not provided here.