About this role
Job title: Staff Software Engineer - Distributed Data Systems
About the Role As a Staff Software Engineer on the Runtime team, you will help build the next generation distributed data storage and processing systems that power large-scale data workloads. You will work on technologies that can outperform specialized SQL engines while providing expressive programming abstractions for ETL, data science, and analytics.
What You'll Do
- Design, implement, and maintain distributed data storage and processing components at Databricks' scale
- Collaborate with engineering and product teams to deliver high-performance, reliable systems across cloud backends (S3, Azure Blob, etc.)
- Improve query planning, execution, and reliability for big data workloads
- Contribute to performance engineering initiatives such as building next-generation optimizers and execution engines
- Mentor more junior engineers and drive architecture and design decisions across projects like Apache Spark, Delta Lake, and Delta Pipelines
- Engage with initiatives in Data Plane Storage and Delta Pipelines to orchestration of large numbers of data pipelines
What We're Looking For
- BS in Computer Science or related field; MS or PhD in databases, distributed systems is optional
- 8+ years of production-level experience in either Java, Scala or C++
- Strong foundation in algorithms and data structures and their real-world use cases
- Experience with distributed systems, databases, and big data systems (Apache Spark, Hadoop)
- Comfortable working towards a multi-year vision with incremental deliverables and a focus on delivering customer value and impact
Nice to Have
- MS or PhD in databases or distributed systems (optional)
Compensation & Benefits
- Local Pay Range: $192,000 — $260,000 USD
- The total compensation package may include annual performance bonus, equity, and the benefits listed above