About this role
Job title: Platform Engineer
About the Role A platform-focused software engineer who builds the cloud-native foundation and automation that powers product teams. You design systems, not tickets, and you orchestrate agents to reduce toil while keeping your hands on architecture and judgment calls.
What You'll Do
- Build the platform. Design and operate the cloud-native foundation on GCP — compute, networking, observability, and the golden paths other teams build on.
- Automate everything that shouldn't be manual. Infrastructure-as-code, self-service pipelines, auto-remediation. If a human repeats it, you write the code — or point an agent at it — to stop it.
- Orchestrate agents for ops. Use AI to draft IaC, generate tests, triage incidents, and summarize signal from noise. You direct, review, and harden — agents do the grunt work.
- Own CI/CD end to end. Fast, reliable, reproducible delivery. Build pipelines that let engineers ship multiple times a day with confidence.
- Engineer reliability in. SLOs, monitoring, alerting that means something, and incident response that gets calmer every quarter. Make on-call boring.
- Secure the path. Bake security and compliance into the platform so doing the right thing is the easy thing — secrets management, least privilege, supply-chain hygiene.
- Support AI workloads. Build the infrastructure that runs LLM and agentic systems in production — scaling, cost control, and the deployment patterns AI apps actually need.
What We're Looking For
- A real programming language, deep — Python and/or Go preferred. You write production code, not just YAML.
- GCP expertise: Cloud Run, GKE, Cloud Functions, networking, IAM. Designed, deployed, and operated at scale.
- IaC & CI/CD: Terraform, Cloud Build, GitHub Actions, or equivalent. Reproducible, version-controlled, peer-reviewed infrastructure.
- Containers & orchestration: Docker and Kubernetes in production, including the parts that break.
- Observability: metrics, logs, tracing. You can find the needle, not just collect haystacks.
- AI-assisted development: fluent with modern AI coding tools and agent workflows, and clear-eyed about where they help and where they don't.
- AI workload experience (bonus): running LLM/agentic systems, GPUs, or model-serving infrastructure in production.
- Reliability & security fundamentals: SLOs, incident response, and security-by-default in pipelines.
- 3+ years in DevOps, SRE, or platform engineering with real software engineering experience.
- 2+ years on cloud-native platforms (GCP preferred) operating production systems.
- BS in Computer Science, Electrical Engineering, or related field; MS a plus but not a substitute for shipping software.
- Demonstrable automation you've built — show the toil you've killed, not just tools you've read about.
Nice to Have
- Experience with AI workloads in production and related tooling.
- Familiarity with GPU-based workloads and model-serving infra.
- Security-focused mindset in pipelines and deployments.
Compensation & Benefits
- Immediate medical, dental, vision and prescription drug coverage.
- Flexible family care days, paid parental leave, new parent ramp-up programs, subsidized back-up child care and more.
- Family building benefits including adoption and surrogacy expense reimbursement, fertility treatments, and more.
- Vehicle discount program for employees and family members and management leases.
- Tuition assistance.
- Established and active employee resource groups.
- Paid time off for individual and team community service.
- A generous schedule of paid holidays, including the week between Christmas and New Year’s Day.
- Paid time off and the option to purchase additional vacation time.