About this role
About the Role
Join an early-stage team building infrastructure that makes production Kubernetes easier to deploy and manage for demanding AI workloads. You will help shape the platform layer and its multi-cluster capabilities, working across systems engineering, automation, and cloud-native infrastructure.
What You'll Do
-
Build platform features in Go, including custom Kubernetes operators and controllers.
-
Design GitOps workflows with ArgoCD to streamline continuous deployment.
-
Develop infrastructure-as-code patterns with Terraform and Helm to provision and manage clusters.
-
Work with distributed storage technologies such as Ceph and WEKA for scalable cluster storage.
-
Create observability systems with Prometheus and Grafana to surface cluster health and performance.
-
Build and optimize container networking with Cilium for security and observability.
-
Design federated Kubernetes architectures for multi-cluster management.
-
Develop automation that reduces operational work for teams running production workloads.
What We're Looking For
-
At least 2 years of software development experience, with approximately 5 years preferred.
-
Hands-on experience writing Go for Kubernetes environments and managing production-scale clusters.
-
Experience with distributed storage such as Ceph or WEKA, federated Kubernetes, and multi-cluster management.
-
Experience with GitOps and ArgoCD, plus infrastructure as code using Terraform, Helm, or comparable tools.
-
Familiarity with cloud platforms, managed Kubernetes, container networking, or service mesh technologies.
-
Experience with Prometheus and Grafana is useful, as are contributions to open-source infrastructure projects.
Compensation & Benefits
Annual salary of $180,000 to $210,000 USD. Visa sponsorship is available.
Location
This is a full-time, on-site role in San Francisco, California, United States.