About this role
Manager, Site Reliability Engineering (SRE)
Location: Remote (US-based candidates only)
Position Overview
We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.
This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.
Key Responsibilities
-
Team Leadership & Development
-
Lead, mentor, and develop a US-based team of Site Reliability Engineers.
-
Conduct regular 1:1s, performance reviews, and career development discussions.
-
Own hiring, onboarding, and retention efforts as the team scales.
-
Foster a culture of ownership, blameless postmortems, and continuous improvement.
-
Operational Excellence & Incident Management
-
Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
-
Serve as an escalation point and incident commander for major production incidents.
-
Drive problem management and root cause analysis processes.
-
Carry PagerDuty on-call escalation responsibilities for critical issues.
-
Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
-
Reliability & Automation
-
Improve system reliability, observability, and resilience using Datadog and related tooling.
-
Drive automation, self-healing capabilities, and runbook maturity.
-
Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
-
Contribute hands-on to tooling, automation, and technical reviews as needed.
-
Collaboration & Global Alignment
-
Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
-
Represent the US SRE organization in cross-functional planning and operational reviews.
-
Communicate effectively with both technical and non-technical stakeholders.
-
Documentation & Compliance
-
Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
-
Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
-
Required Qualifications
-
5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
-
1–2+ years of experience leading, mentoring, or managing engineers.
-
Demonstrated success operating in a player-coach leadership model.
-
Strong hands-on experience with production incident management and escalation processes.
-
Proficiency with Datadog or similar observability platforms.
-
Hands-on experience with Kubernetes and Docker in production environments.
-
Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
-
Experience with Helm, CI/CD pipelines, and deployment automation.
-
Working knowledge of ITIL processes and Agile methodologies.
-
Experience working with SQL, MySQL, or NoSQL databases.
-
Excellent communication and stakeholder management skills.
-
Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.
-
Preferred Qualifications
-
Experience with cloud platforms such as AWS, Azure, or GCP.
-
Experience building or scaling SRE teams and on-call programs.
-
Experience defining and managing SLOs, SLIs, and error budgets.
-
Prior experience in the healthcare or fintech industry.
-
Knowledge of security and compliance frameworks relevant to regulated environments.
-
Why Join NationsBenefits?
-
Competitive compensation and comprehensive benefits.
-
Unlimited PTO.
-
Fully remote work environment (US-based).
-
Opportunity to lead and grow a high-impact SRE organization.
-
Exposure to modern cloud-native technologies and large-scale reliability challenges.
-
Collaborative culture focused on innovation, learning, and continuous improvement.
-
Meaningful work that directly impacts healthcare technology and millions of members.
-
Ideal Candidate
-
We are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.
-
NationsBenefits is an Equal Opportunity Employer.