About this role
Job title: Site Reliability Engineer (SRE)
About the Role Retool's Core Infrastructure team owns the systems that make Retool run for enterprise customers, including Retool Cloud, managed single-tenant environments, BYOC, Kubernetes and Helm deployments, and the migration paths between them. As an SRE, you will own reliability across these deployment paths and help customers upgrade and operate at scale with safety and clarity.
What You'll Do
- Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
- Build automation to turn today’s manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
- Improve observability for Retool Cloud, self-hosted customers, and internal operators, turning health signals into clear status, likely causes, and recommended actions.
- Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current.
- Help move customers from legacy or less-supported deployment models toward supported paths such as Retool’s official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows repeatable for customers, Support, and TAMs.
- Partner with product engineers on infrastructure requirements for new Retool products, especially when they introduce new dependencies.
- Lead through ambiguity, make careful risk calls, and communicate clearly while things are moving quickly.
- Write the docs, runbooks, design notes, and migration guides that make complex systems understandable to other engineers and to customers.
What We're Looking For
- Infrastructure fundamentals: Deep experience operating production infrastructure in AWS; strong Kubernetes fundamentals; real Terraform or infrastructure-as-code experience; good operational judgment around Postgres.
- Reliability and automation: Experience building or operating observability systems; a bias toward automation.
- Programming ability: Proficiency in Go, Python, TypeScript, Java, or Ruby.
- Customer and team judgment: Clear written communication; comfort working directly with customer-facing teams and, when useful, customers themselves.
Nice to Have
- Experience debugging customer environments and working in environments with ambiguity and rapid change.
- Ambitious, curious, energetic, and careful with the details.