About this role
About the Role Design, build, and maintain data pipelines and ETL with Databricks and Apache Spark. Implement data Lakehouse architecture, manage data ingestion from multiple sources, and ensure data quality, governance, and security. Collaborate with data scientists and analysts to enable advanced analytics and machine learning workloads. Monitor and troubleshoot Databricks clusters, jobs, and workflows, and integrate Databricks with cloud services (AWS, Azure, or GCP). Document processes, standards, and best practices for data engineering. What You'll Do
- Design, build, and maintain data pipelines and ETL using Databricks and Apache Spark.
- Optimize data workflows for performance, scalability, and cost efficiency.
- Implement data Lakehouse architecture and manage ingestion from multiple sources.
- Collaborate with data scientists and analysts to enable advanced analytics and ML workloads.
- Ensure data quality, governance, and security across all data assets.
- Monitor and troubleshoot Databricks clusters, jobs, and workflows.
- Integrate Databricks with cloud services (AWS, Azure, or GCP) and other enterprise systems.
- Document processes, standards, and best practices for data engineering. What We're Looking For
- Experience with Delta Lake and Lakehouse architecture.
- Knowledge of Azure Data Factory (ADF) for orchestration.
- Exposure to CI/CD pipelines for data engineering.
- Experience with SQL and Spark SQL.
- Azure certifications related to data engineering.