About this role
About the Role Alloy is seeking an Infrastructure & Security Lead Site Reliability Engineer to design and scale automated infrastructure for a large and growing footprint. As part of our Infrastructure Team, you’ll help reduce toil, improve reliability, and enable safe self-service for other engineers. This role blends hands-on engineering with architectural decisions to ensure scalable, secure operations across Kubernetes, databases, and services.
What You'll Do
- Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)
- Reduce operational toil by turning manual processes into reliable, repeatable workflows
- Build internal tooling and platforms that enable safe self-service changes for other engineers
- Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)
- Implement and evolve systems for deploying and running applications in Kubernetes
- Contribute to architecture decisions across infrastructure, reliability, and security
- Write and review production-quality code
- Participate in on-call rotations—but focus on building systems that prevent incidents, not just respond to them
What We're Looking For
- 10+ years of experience in infrastructure, SRE, or software engineering roles
- Strong software engineering skills—you build systems, not just scripts
- Experience managing production infrastructure at scale (cloud + containerized systems)
- Experience with Infrastructure as Code (e.g., Terraform)
- Experience running and troubleshooting distributed systems (Docker/Kubernetes)
- Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)
- Proficiency in at least one programming language (Python, Go, JavaScript, etc.)
- Experience participating in on-call rotations and improving systems based on incidents
- Strong communication and collaboration skills
Nice to Have
- Experience running Kubernetes in production at scale
- Deep familiarity with AWS
- Experience building internal platforms or developer tooling
- Background in distributed systems or large-scale data systems
Compensation & Benefits
- Salary range: $179,000 to $226,000 per year
- Benefits and Perks: Unlimited PTO and flexible work policy; Employee stock options; Medical, dental, vision plans with HSA and FSA options; 401k with 100% match up to 4% of annual employee compensation; Eligible new parents receive 16 weeks of paid parental leave; Home office stipend for new employees; Annual Learning & Development annual stipend; Well-being benefits include access to ClassPass, OneMedical, UrbanSitter, and Spring Health; Hybrid work environment: Tuesdays–Thursdays from HQ in Union Square, Manhattan; Monday/Friday remote; Lunch catered from local restaurants; Frequent employee-organized cultural events; Many employees Zoom into work from home on Mon/Fri.