About this role
Lead Data Engineer with 7-10 years of experience to build and optimize scalable ETL/ELT pipelines using Azure Databricks, PySpark, and Delta Lake. The role involves working across scrum teams to develop data solutions, ensure data governance with Unity Catalog, and support real-time and batch processing in the insurance and healthcare domain.
Collaborate with multiple scrum teams to deliver quality database programming requirements for each sprint. Leverage Azure cloud platforms (Azure Databricks, Advanced Python Programming, Azure SQL, Data Factory, Data Lake, ADLS Gen2) to build scalable ETL/ELT pipelines. Create and deploy scalable ETL/ELT pipelines in Azure Databricks using PySpark and SQL. Create Delta Lake tables with ACID transactions and schema evolution to support real-time and batch processing. Use Unity Catalog for centralized data governance, access control, and data lineage tracking. Independently analyze, diagnose, and resolve issues in real time; provide end-to-end problem resolution. Develop unit tests to enable automated testing. Apply SOLID development principles to maintain data integrity and cohesion. Collaborate with product owners and business stakeholders to determine and satisfy needs. Demonstrate ownership, accountability, and impact on the company's success. Exhibit critical thinking, problem-solving, teamwork, time-management, and strong interpersonal and communication skills.
7-10 years of experience as a Lead Data Engineer. Self-driven with minimal supervision. Proven experience with T-SQL programming, Azure Databricks, Spark (PySpark/Scala), Delta Lake, Unity Catalog, and ADLS Gen2. Microsoft TFS, Visual Studio, DevOps exposure. Experience with cloud platforms such as Azure. Analytical, problem-solving mindset.
Nice to Have: Healthcare domain knowledge.