Talent Apply
Log in
All jobs
D

Incident Manager

databricks
Remote (US)
RemoteUSD 103,900 - 145,525 / year

About this role

Job title: Incident Manager

About the Role As an Incident Manager, you will lead Databricks’ most critical production incidents, providing clear, timely communication to customers, executives, and engineers. You’ll serve as incident commander and reliability engineer, orchestrating multi-team responses and driving real-time status updates to ensure technical resilience and stakeholder confidence.

What You'll Do

  • Lead critical incidents across Databricks’ cloud-based services, coordinating multi-disciplinary response to rapidly mitigate impact and restore operations.
  • Drive technical root cause analysis and reliability improvements with engineering teams; trace and document underlying causes across distributed systems and data stores; summarize learnings and ensure action items are followed through.
  • Own communications during incidents; deliver frequent, high-quality updates to internal stakeholders and compose customer-facing notifications that are accurate, timely, and empathetic.
  • Develop and maintain incident playbooks and communication templates to ensure consistent, timely updates.
  • Mentor and train peers in incident communication and technical response to raise overall incident response quality.

What We're Looking For

  • 5+ years of experience in incident management, site reliability engineering, or production operations supporting large-scale, cloud-native systems.
  • Proven ability to lead high-severity incidents, identify impact, isolate fault domains, and manage multi-team response efforts.
  • Strong understanding of cloud infrastructure (AWS, Azure, or GCP) — including compute, networking, storage, and observability components.
  • Deep expertise in log analysis and debugging with tools such as Datadog, Elasticsearch, Splunk, Cloud Logging, or OpenTelemetry.
  • Hands-on experience with observability systems (metrics, logging, tracing) using Prometheus, Grafana, OpenTelemetry, etc.
  • Proficiency in at least one major programming or scripting language (Python, Go, or Bash) for automating diagnostics, data collection, or analysis.
  • Experience developing and maintaining incident playbooks and communication templates to ensure consistent, timely updates.
  • Excellent contextual interpretation and writing skills; ability to communicate to technical and business audiences.
  • BS, MS, or related degree in Computer Science, Computer Engineering, or related field.

Compensation & Benefits

  • Zone 3 Pay Range: $103,900 - $145,525 USD annually.
  • The total compensation package may include annual performance bonus, equity, and benefits. For more information, visit the location page.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →