About this role
Lead - Cloud and Data Platform Operations
Bangalore, Karnataka, India
Job Description
About Chubb
Chubb is a world leader in insurance. With operations in 54 countries and territories, Chubb provides commercial and personal property and casualty insurance, personal accident and supplemental health insurance, reinsurance and life insurance to a diverse group of clients. The company is defined by its extensive product and service offerings, broad distribution capabilities, exceptional financial strength and local operations globally. Parent company Chubb Limited is listed on the New York Stock Exchange (NYSE: CB) and is a component of the S&P 500 index. Chubb employs approximately 40,000 people worldwide. Additional information can be found at: www.chubb.com.
About Chubb India
At Chubb India, we are on an exciting journey of digital transformation driven by a commitment to engineering excellence and analytics. We are proud to share that we have been officially certified as a Great Place to Work® for the third consecutive year, a reflection of the culture at Chubb where we believe in fostering an environment where everyone can thrive, innovate, and grow
With a team of over 2500 talented professionals, we encourage a start-up mindset that promotes collaboration, diverse perspectives, and a solution-driven attitude. We are dedicated to building expertise in engineering, analytics, and automation, empowering our teams to excel in a dynamic digital landscape.
We offer an environment where you will be part of an organization that is dedicated to solving real-world challenges in the insurance industry. Together, we will work to shape the future through innovation and continuous learning.
Position Details
Job Title: Cloud and Data Platform Operations Lead
Function/Department: Technology
Employment Type: Full-time
Location: Hyderabad/Bangalore
Position Overview
We are seeking an experienced operations leader to head the Cloud and Data Platform Operations function out of Chubb's engineering centre in India. This is an Operations Leadership role — the successful candidate will own the 24/7 availability, reliability, and performance of Chubb's enterprise cloud infrastructure (AWS, Azure, GCP) and data platforms hosted on these hyperscalers.
This role leads a globally distributed, follow-the-sun operations team spanning APAC, EMEA, and the Americas, ensuring continuous 24/7 coverage with no gaps across time zones. You will be the senior escalation point for critical production incidents at any hour, driving rapid resolution and systematic prevention.
Alongside operational accountability, you will champion engineering discipline within the team — driving automation, Infrastructure-as-Code (IaC), DevSecOps practices, and data platform optimization to reduce toil, improve resilience, and support safe, frequent delivery.
Candidates should note: this role carries material out-of-hours responsibility. If you are seeking a primarily engineering or architecture role without significant operational accountability, this position is not the right fit.
Key Responsibilities
-
- 24/7 Operations & Service Reliability
-
Own end-to-end production operations for cloud infrastructure and data platforms, ensuring systems meet availability, performance, and SLA commitments around the clock.
-
Serve as the senior escalation point for critical incidents — leading response, coordination, and communication across technical teams and executive stakeholders, regardless of time zone.
-
Establish and maintain on-call rotas, escalation procedures, and runbooks across the globally distributed team to ensure no single point of failure in operational coverage.
-
Drive post-incident reviews (PIRs) and root cause analysis (RCA), translating findings into preventative engineering improvements and updated operational playbooks.
-
Define and track operational KPIs — MTTR, MTTD, change failure rate, system availability — and report these clearly and regularly to senior leadership.
-
- Follow-the-Sun Team Leadership
-
Lead, manage, and develop a high-performing, globally distributed operations team across multiple time zones (APAC, EMEA, Americas), ensuring seamless handoffs and continuous 24/7 operational coverage.
-
Build and enforce structured shift handover processes so that context, active incidents, and work-in-progress are reliably transferred between regional teams without information loss.
-
Foster a strong operational culture: disciplined on-call practice, blameless post-mortems, knowledge sharing, and continuous improvement of operational standards.
-
Mentor team members across regions, balancing local autonomy with consistent global standards and escalation paths.
-
Partner across cross-functional engineering, security, and business groups to align operational delivery with executive timelines and organizational priorities.
-
- Incident, Problem & Change Management (ITSM)
-
Own the full ITIL/ITSM lifecycle using ServiceNow: Incident, Problem, Change, and Request Management for all cloud and data platform services.
-
Lead Major Incident Management for P1/P2 disruptions — coordinating response bridges, managing stakeholder communications, and driving resolution with urgency and clarity.
-
Drive the Problem Management process: identify recurring incidents, prioritize root cause elimination, and track remediation to closure.
-
Govern the Change Management process, ensuring all production changes are reviewed, risk-assessed, and scheduled to minimize operational disruption.
-
- Multi-Cloud Infrastructure & Platform Engineering
-
Design, build, and operate secure, highly available, and scalable cloud infrastructure across AWS, Azure, and GCP — with a focus on operational manageability and resilience.
-
Deliver automated, reusable platform components and self-service templates (modular IaC via Terraform/Bicep) to reduce manual operational toil and improve consistency.
-
Implement robust CI/CD pipelines with embedded security and quality gates (DevSecOps), supporting safe and frequent production deployments.
-
Coordinate vendor relationships, manage cloud costs under a FinOps framework, and ensure platform components meet operational standards before production deployment.
-
- Data Platform Operations & Spend Optimization
-
Operate and maintain scalable data pipelines, ETL/ELT processes, and data models supporting underwriting algorithms and analytics workloads — with a focus on production reliability and data quality.
-
Manage and optimize Azure Data Lake Storage (ADLS), Databricks, Snowflake, and Azure Synapse Analytics environments: cluster sizing, job scheduling, performance tuning, and cost governance.
-
Operate and maintain Astronomer (managed Apache Airflow) for pipeline orchestration — ensuring DAG reliability, SLA adherence, and efficient resource utilization across workflow runs.
-
Enforce data lifecycle policies, monitor consumption, and maintain strict production readiness standards for Lakehouse and data warehouse platforms.
-
Ensure data quality, integrity, and compliance controls are embedded into operational processes, not treated as afterthoughts.
-
- Stakeholder Management & Operational Reporting
-
Translate operational performance — availability metrics, incident trends, cost data, platform health — into clear, actionable insights for non-technical executives and corporate underwriting leadership.
-
Manage stakeholder expectations proactively: communicate planned maintenance, emerging risks, and incident impact with transparency and appropriate urgency.
-
Act as the operational voice in delivery planning, ensuring engineering teams understand production constraints, operational readiness requirements, and go-live criteria before releases.
Qualifications
- Education: Bachelor's degree in IT, Computer Science, Engineering, or a related discipline.
- Experience: 15–20 years of progressive, hands-on experience across cloud operations, platform engineering, infrastructure, or data engineering — with a significant portion in production operations roles.
- 24/7 Operations
Experience
- Demonstrated experience owning or leading production operations in a 24/7, always-on environment. Comfort with being an escalation point for critical incidents outside business hours is essential.
- Follow-the-Sun Leadership: Proven track record managing and coordinating a geographically distributed team operating across multiple time zones, with structured handover processes and continuous coverage models.
- ITSM & Incident Management: Thorough, hands-on knowledge of ITIL/ITSM frameworks. Functional fluency with ServiceNow for Incident, Problem, Change, and Request management. Experience leading Major Incident bridges.
- Multi-Cloud Technical Stack: Deep proficiency in native services across AWS, Azure, and GCP, including container orchestration, networking, identity, and multi-cloud landing zone design.
- DevSec