About this role
Job title: Incident Manager
About the Role As an Incident Manager, you will lead Databricks’ most critical production incidents, providing clear, timely communication to customers, executives, and engineers. You’ll serve as incident commander and reliability engineer, orchestrating multi-team responses and driving real-time status updates to ensure technical resilience and stakeholder confidence.
What You'll Do
- Lead critical incidents across Databricks’ cloud-based services, coordinating multi-disciplinary response to rapidly mitigate impact and restore operations.
- Drive technical root cause analysis and reliability improvements with engineering teams; trace and document underlying causes across distributed systems and data stores; summarize learnings and ensure action items are followed through.
- Own communications during incidents; deliver frequent, high-quality updates to internal stakeholders and compose customer-facing notifications that are accurate, timely, and empathetic.
- Develop and maintain incident playbooks and communication templates to ensure consistent, timely updates.
- Mentor and train peers in incident communication and technical response to raise overall incident response quality.
What We're Looking For
- 5+ years of experience in incident management, site reliability engineering, or production operations supporting large-scale, cloud-native systems.
- Proven ability to lead high-severity incidents, identify impact, isolate fault domains, and manage multi-team response efforts.
- Strong understanding of cloud infrastructure (AWS, Azure, or GCP) — including compute, networking, storage, and observability components.
- Deep expertise in log analysis and debugging with tools such as Datadog, Elasticsearch, Splunk, Cloud Logging, or OpenTelemetry.
- Hands-on experience with observability systems (metrics, logging, tracing) using Prometheus, Grafana, OpenTelemetry, etc.
- Proficiency in at least one major programming or scripting language (Python, Go, or Bash) for automating diagnostics, data collection, or analysis.
- Experience developing and maintaining incident playbooks and communication templates to ensure consistent, timely updates.
- Excellent contextual interpretation and writing skills; ability to communicate to technical and business audiences.
- BS, MS, or related degree in Computer Science, Computer Engineering, or related field.
Compensation & Benefits
- Zone 3 Pay Range: $103,900 - $145,525 USD annually.
- The total compensation package may include annual performance bonus, equity, and benefits. For more information, visit the location page.