Talent Apply
Log in
All jobs
A

Site Reliability Engineer III

Amgen
Hyderabad, Telangana, India
On-site

About this role

About the Role

The GCF5 Site Reliability Engineer is the senior technical leader for the HPC Enablement pillar. They define and socialize operational standards and patterns, lead multi-team delivery, mentor GCF4 engineers, and translate researcher needs into scalable compute enablement designs. They own pillar-level reliability, performance, cost efficiency, and SLA/SLO outcomes, and influence cross-team engineering quality. This role reports to the GCF7 leader and partners closely with peer GCF5 domain leads across SCIP to ensure cohesive, scalable platform evolution. What You'll Do

  • Own the compute reliability and enablement roadmap within SCIP.

  • Define onboarding playbooks and golden paths for HPC workloads.

  • Establish containerization and reproducible runtime standards.

  • Optimize scheduler configuration and resource allocation policies.

  • Conduct workload profiling and performance tuning.

  • Define and manage SLOs, reliability standards, and operational guardrails.

  • Lead incident response and reduce recurring failures.

  • Mentor engineers and elevate reliability practices.

  • Partner with scientific teams to translate compute requirements into scalable infrastructure patterns. What We're Looking For

  • Deep expertise in HPC Enablement (HPC) with evidence of standard‑setting and reuse.

  • Systems design at scale (HPC); performance, security, and observability fundamentals.

  • Product/engineering thinking: road mapping, prioritization, and outcome‑oriented delivery.

  • Stakeholder influence across science, engineering, and governance forums; crisp written/verbal communication.

  • Basic Qualifications:

  • BS+8 / MS+6 / PhD in CS/Engineering/Data disciplines.

  • Demonstrated production delivery experience in HPC at scale.

  • Demonstrated literacy in a relevant scientific domain (e.g., biology, chemistry, therapeutic discovery).

  • Preferred Qualifications:

  • Depth in HPC Enablement (HPC).

  • Kubernetes and continuous integration/continuous deployment (CI/CD). Nice to Have

  • Depth in HPC Enablement (HPC).

  • Kubernetes and CI/CD. Compensation & Benefits Not specified in the posting.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →