About this role
Job title: HPC Data Center Operational Lead
About the Role Jump's HPC Data Center Operational Lead is responsible for the reliability, standards, and day-to-day operation of Jump's HPC data centers. You will lead multiple site teams, drive preventative maintenance, and leverage AI-driven insights to maximize uptime in a fast-paced, research-driven environment.
What You'll Do
- Lead and manage data center site leads and their teams across multiple HPC facilities; site leads report directly to this role.
- Recruit, mentor, and develop team members while conducting performance reviews and building a culture of operational rigor.
- Direct onsite contractors with clear scope and validation of completed work.
- Develop, document, and enforce operational standards for Jump's HPC data centers covering power, cooling, cabling, and hardware lifecycle.
- Design and own the preventative maintenance program, including scheduled inspections, component replacements, and firmware/capacity reviews to minimize unplanned downtime.
- Drive continuous improvement of operational processes and pursue automation—including AI-driven approaches—to reduce manual effort and human error.
- Serve as the subject matter authority on HPC data center power distribution, power striping strategies, and failover/redundancy configurations.
- Own expertise across air cooling, liquid cooling (direct-to-chip, rear-door, CDU-based), and hybrid cooling architectures.
- Maintain deep knowledge of environmental monitoring and controls (temperature, humidity, airflow, leak detection) and ensure systems remain within design parameters.
- Own the HPC data center monitoring strategy end-to-end: define what is monitored, set alerting thresholds, and ensure comprehensive visibility into facility and hardware health.
- Leverage AI tools to analyze telemetry data, identify failure patterns, predict potential issues, and accelerate root cause analysis during incidents.
- Lead critical incident response and drive root cause analysis and corrective actions to prevent recurrence.
- Establish and track operational KPIs including availability, mean time to repair, and efficiency metrics.
- Maintain deep, hands-on knowledge of server hardware architectures including multi-socket platforms, GPU/accelerator configurations, memory subsystems, NVMe/storage controllers, BMC/IPMI management, and firmware lifecycle.
- Maintain deep, hands-on knowledge of network switch hardware including line cards, optics/transceivers, switch fabrics, and platform-specific diagnostics for Arista and Cisco platforms.
- Evaluate new hardware platforms, drive hardware qualification and acceptance testing, and provide informed recommendations on hardware selection.
- Own the overall hardware break-fix function across all HPC sites, ensuring rapid diagnosis and resolution for servers, GPUs, network equipment, storage, and facility infrastructure.
- Diagnose complex hardware failures at the component level—CPUs, DIMMs, GPUs, NICs, PSUs, fans, drives, switch line cards, and optics—and direct the team to resolve efficiently.
- Establish escalation paths, SLA targets, and reporting for hardware failures.
- Own inventory processes and spares tracking across all HPC facilities, ensuring critical spares are stocked, tracked, and replenished to meet availability targets.
- Maintain accurate asset records for all serialized and consumable inventory.
- Collaborate with planning and vendor teams to ensure capacity and resilience across facilities.
What We're Looking For
- A seasoned operations leader with experience running data center operations across multiple HPC facilities.
- Deep expertise in HPC data center power distribution, power striping, and failover/redundancy.
- Strong knowledge of air cooling, liquid cooling (direct-to-chip, rear-door, CDU-based), and hybrid cooling architectures.
- Proficiency in environmental monitoring and controls; ability to keep systems within design parameters.
- Experience driving monitoring strategy, defining metrics, and using AI/telemetry for predictive maintenance.
- Excellent incident response and root cause analysis skills.
- Hands-on server hardware knowledge (multi-socket CPUs, GPUs, NVMe storage, firmware lifecycle) and network hardware (Arista, Cisco).
- Experience with hardware qualification, testing, and making informed hardware recommendations.
- Demonstrated ability to lead hardware break-fix, diagnose component-level failures, and manage escalation and SLAs.
- Strong inventory and spares management, asset tracking, and vendor planning.
- Willingness to be on-site 5 days/week in Chicago or New York and travel as needed.
Nice to Have
- Experience applying AI-driven automation to data center operations.
- Familiarity with HPC workloads and associated tooling.
- Experience with BMS/IPMI and advanced environmental controls.
- Prior experience in vendor management and planning across multiple facilities.