About this role
Optum Insight Engineering
The Optum Insight Engineering team is seeking a Lead AI Ops Cloud Engineer with Site Reliability Engineering (SRE) experience to design, build, and scale modern AI Ops solutions across payer and provider transformation initiatives. This is a hands-on technical leadership role requiring deep involvement in architecture and engineering while leading globally distributed teams. The role focuses on delivering, Agentic AI, and Retrieval Augmented Generation (RAG) solutions with enterprise-grade reliability, security, and responsible AI practices.
You'll enjoy the flexibility to work remotely from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities
- Lead end-to-end design and implementation of AI Ops solutions from concept through production with an emphasis on responsible AI practices
- Contribute to the improvement of our SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments
- Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform & GitHub Actions while adhering to best practices
- Define and own solution architecture for RAG pipelines, agentic workflows, tool calling, orchestration, and conversational context management
- Design, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions
- Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers
- Design, code, test, and operate software using Python & Node.js
- Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiency
You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications
- 8+ years of overall software engineering and Site Reliability Engineering (SRE) experience in a public cloud environment (GCP, AWS, Azure)
- 3+ years of experience of demonstrated hands-on experience with Python and Terraform based development
- 2+ years delivering AI/ML or Generativ