About this role
Job title: Senior Staff Backline Engineer - Data & AI
About the Role Senior Staff Backline Engineer - Data & AI at Databricks serves as the critical bridge between Frontline Support and Engineering to resolve high-priority architectural failures and complex system anomalies. You will lead deep-dive troubleshooting, root-cause analysis, and architectural optimization to improve stability, throughput, and long-term supportability of the platform.
What You'll Do
- Conduct deep-dive forensics into Spark core internals and the broader Databricks Data and AI ecosystem to resolve high-priority architectural failures and complex system anomalies.
- Perform advanced code-level root-cause analysis and resource profiling to identify systemic issues, ensuring stability of high-scale production workloads.
- Optimize architectural performance by refining execution parameters and enforcing best-practice strategies to maximize resource efficiency and throughput.
- Analyze global issue trends and partner with Product Engineering to influence the product roadmap and drive initiatives that enhance long-term supportability.
- Develop reproduction frameworks, automated workflows, and AI-driven diagnostic tools that translate backline findings into standardized resolution paths to empower and scale the organization.
What We're Looking For
- 10+ years of relevant experience, including deep expertise in one of three specialized tracks, with proven experience managing both customers and technical stakeholders.
- Data Engineering Track: Expertise in large-scale big data solutions (Spark, Delta Lake, Hive); strong troubleshooting, diagnosing performance issues, and identifying root causes; solid hands-on programming in Python, SQL, or Scala.
- Product Supportability Track: Deep understanding of distributed system internals; ability to perform code-level root-cause analysis and profiling (Java, Scala, or Python); proven record of contributing to bug fixes and mentoring engineers.
- AI Track: Experience with large-scale ML and generative AI systems (LLMs, agent-driven workflows); knowledge of model training, evaluation, deployment in distributed environments; experience governing ML lifecycle and operationalization; diagnosing and optimizing distributed ML workloads for performance and scalability.
Compensation & Benefits
- Local Pay Range $170.40—$255.60 USD per hour.
- Total compensation may include annual performance bonus, equity, and benefits. Location-dependent.