Talent Apply
Log in
All jobs
I

Principal Software Engineer – SRE, Incident Response & Operational Excellence

Intuit
San Diego, California; Mountain View, California
On-site

About this role

Principal Software Engineer – SRE, Incident Response & Operational Excellence

Category Software Engineering

Location

San Diego, California; Mountain View, California

Job ID 24502

Apply Now

Share Job Save JobSaved Job

Company Overview

Intuit is the global financial technology platform that powers prosperity for the people and communities we serve. With tens of millions of customers worldwide using products such as TurboTax, Credit Karma, QuickBooks, and Mailchimp, we believe that everyone should have the opportunity to prosper. We never stop working to find new, innovative ways to make that possible.

Job Overview

Intuit is the global financial technology platform that powers prosperity for the people and communities we serve. Our public-facing products — TurboTax, Credit Karma, QuickBooks, and Mailchimp — serve approximately 100 million customers who depend on us at the moments that matter most: filing a return before the deadline, running payroll, checking their credit, or sending a campaign. When something breaks, minutes matter, and how quickly and how well we recover is a direct expression of our customer promise.

We are hiring Principal Software Engineers to serve as Technical Duty Officers (TDOs) — the leaders who take command of Intuit’s most critical incidents. During our most critical incidents you hold ultimate authority on decisions and enforcement of our incident principles, partnering with the Business Unit Driver and the Incident Operations Center (IOC) to direct recovery, actions, and communications on the bridge. Reporting to the Director of Cloud Engineering & Operations, this is a dedicated, full-time role: when you are not on the bridge, your mandate is to make the next incident shorter, rarer, and better understood — partnering across every business unit on effective observability, clarity of customer workflows and their real-time impact signals, high-quality runbooks, and validated escalation and notification readiness.

This role is equal parts operator, engineer, and leader. You bring the calm, decisive command presence to lead senior leaders through high-stakes ambiguity, the deep systems intuition to cut to the most likely fault and the fastest safe path to recovery, and the engineering skills to implement the improvements you identify rather than hand them off. As a Principal engineer, you will not just execute this work manually but direct, orchestrate, and govern AI agents to accelerate detection, triage, impact assessment, and post-incident learning, accountable for the quality and reliability of everything you and those agents produce.

Responsibilities

  • Lead Intuit’s most critical incidents, serve as the “Technical Duty Officer” role on all significant incidents, holding ultimate authority on decisions and enforcement of incident principles; direct recovery, actions, and communications on the bridge in partnership with the BU Driver and IOC Leader

  • Drive fast recovery, bias the bridge toward mitigation over diagnosis, make timely and decisive calls on rollbacks, failovers, traffic shifts, and feature disablement, and keep responders focused on the shortest safe path to restoring the customer experience

  • Get the priority right early, challenge and correct incident priority based on customer impact during active response rather than after recovery, and ensure the right leaders, escalations, and communications are engaged at the right time

  • Establish clear customer-impact signals, partner with each business unit to define, instrument, and continuously report real-time customer impact for the highest-value workflows across top revenue-generating products, and drive readiness validation through the company-wide Customer Impact Gameday

  • Raise the observability bar, partner with platform and product teams on SLOs, dashboards, alerting, and tracing so responders can assess the problem and its impact in minutes; feed tool and product-gap findings from live incidents back to platform teams as part of the RCA

  • Own runbook and workflow clarity, ensure critical customer workflows have accurate, exercised runbooks and clear service ownership, and drive the recurring validation of notification channels, escalation policies, bridge creation, and executive notification workflows

  • Make communications actionable, work with the IOC Leader so that executive and customer-facing updates consistently convey customer impact, recovery progress, and next steps with enough clarity for confident decision-making

  • Lead post-incident learning, drive blameless RCAs for the incidents you command, identify systemic and cross-BU remediation, and track high-leverage actions to closure

  • Continuously improve incident and observability capabilities, bring an SRE mindset to the incident-management and observability platforms, implement solutions yourself, and encode triage, impact-assessment, and communication knowledge into agentic and self-service tooling so standards travel with the automation

  • Provide 24x7 coverage, participate in the follow-the-sun TDO rotation across the US and IDC, including off-hours and peak-season (e.g., tax deadline, payroll) readiness windows, with clean shift handoffs

  • Coach and scale incident leadership, contribute to Incident Commander training for Directors, Distinguished Engineers, and Vice Presidents, mentor BU Drivers and responders on the bridge, and raise the incident-leadership bar across the company

Qualifications

  • BS in Computer Science or equivalent work-related experience; 12+ years of engineering experience, including significant time operating large-scale, high-traffic production systems

  • Proven track record commanding large-scale incidents in an enterprise-level or financial-services environment, with the ability to lead senior leaders, make decisive calls with incomplete information, and remain calm and clear under pressure

  • Deep operational excellence and SRE expertise, including SLOs and error budgets, observability (metrics, logs, traces), alerting design, on-call and escalation practices (e.g., PagerDuty), and formal incident-management frameworks and change management

  • Broad, hands-on distributed-systems intuition across the full stack — cloud infrastructure (AWS), Kubernetes and service meshes, networking and edge (CDN, DNS, TLS), data stores, and messaging — with the ability to reason quickly about failure modes, blast radius, and the fastest safe path to recovery

  • Demonstrated ability to translate technical state into crisp customer-impact and business-impact terms for executives and customer-facing teams, in writing and live on a bridge

  • A software engineer who can build, not just operate: proficient in scripting and development (e.g., Python, Golang, Java), Infrastructure as Code, and automation, able to implement or oversee the observability, runbook, and escalation improvements you identify

  • Fluent in AI agent development and orchestration, able to build agentic and self-service tooling (e.g., MCP servers and skills) for triage, impact assessment, and post-incident analysis, and to review AI-generated output to the same quality bar as peer work

  • Experience driving blameless post-incident reviews and systemic remediation across organizational boundaries, influencing Director- and VP-level leaders without direct authority

  • Willingness to participate in a 24x7 follow-the-sun on-call rotation, including off-hours and peak-season coverage

  • Principal-level expectations: align on the technical vision for incident response, observability, and operational readiness across multiple business units; define and drive the company-wide standards, guardrails, and GameDays that measurably reduce time-to-action and time-to-recovery; and be recognized as the authority others turn to during Intuit’s most critical moments

  • Intuit provides a competitive compensation package with a strong pay for performance rewards approach. This position may be eligible for a cash bonus, equity rewards and benefits, in accordance with our applicable plans and programs (see more about our compensation and benefits at Intuit®: Careers | Benefits). Pay offered is based on factors such as job-related knowledge, skills, experience, and work location. To drive ongoing fair pay for employees, Intuit conducts regular comparisons across categories of ethnicity and gender.

The expected base pay range for this position is:

  • San Diego $247,500 - $335,000

  • Mountain View, CA $261,500- $353,500

  • Apply Now

  • Save JobSaved

  • Share Job

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →