About this role
Job title: Senior Principal Data Engineer
Job description
About this role
Role Overview
Data and AI are central to our platform and competitive advantage. We are seeking a Senior Principal Data Engineer to design, build, and operate large‑scale, enterprise data platforms, with deep AI/ML and GenAI expertise—not only enabling AI use cases, but using AI as a first‑class tool to design, build, and optimize data pipelines.
This role is a hands‑on senior engineering position with architectural ownership, operating at the intersection of data engineering, AI/ML systems, and platform modernization.
Key Responsibilities
-
Architect, build, deploy, and operate scalable ETL/ELT data pipelines supporting analytics, ML, and GenAI workloads.
-
Use AI/ML and GenAI to design and build data pipelines, including: AI‑assisted pipeline and schema design; AI‑generated and optimized transformation logic and code; automated test generation, data validation, and performance tuning; AI‑driven documentation and operational runbooks.
-
Apply AI/ML techniques to automate data engineering workflows, including ingestion, schema evolution, data quality checks, anomaly detection, and pipeline optimization.
-
Build and maintain feature engineering pipelines, training datasets, and feedback loops for production ML and GenAI systems.
-
Enable GenAI use cases, including unstructured data ingestion, retrieval pipelines (RAG), and governance‑aware AI data flows.
-
Ensure enterprise standards for reliability, observability, lineage, security, and compliance across all data and AI pipelines.
-
Partner with product, program, analytics, and AI teams in Agile delivery models to translate business needs into scalable AI‑enabled data solutions.
-
Provide senior technical leadership through architecture reviews, mentoring, and setting data & AI engineering standards.
-
Required Qualifications
-
15+ years of experience owning and operating large‑scale data engineering platforms in production.
-
Strong hands‑on expertise in Python and SQL, with deep experience on cloud data platforms (AWS, Azure, or GCP).
-
Proven experience architecting distributed, high‑performance data systems.
-
Demonstrated experience using AI/ML or GenAI to design, build, or optimize data pipelines (not just consuming AI outputs).
-
Strong experience supporting ML/AI systems in production, including feature pipelines, data validation, monitoring, and retraining support.
-
Solid understanding of data modeling, performance tuning, DevOps/CI‑CD, and Agile delivery.
-
Excellent communication skills and ability to operate in fast‑paced, cross‑functional environments.
-
Preferred Skills
-
Experience with GenAI/LLM ecosystems (e.g., RAG, vector data, prompt or model lifecycle support).
-
Experience with microservices, APIs, and large‑scale analytics platforms.
-
Prior experience leading or mentoring senior engineers (player‑coach model).
-
Experience working in financial services or regulated environments.
-
Why This Role
-
This role goes beyond tr