Talent Apply
Log in
All jobs
JT

HPC Data Center Operational Lead

Jump Trading
Chicago, IL or New York, NY
On-site

About this role

Job title: HPC Data Center Operational Lead

About the Role Jump's HPC Data Center Operational Lead is responsible for the reliability, standards, and day-to-day operation of Jump's HPC data centers. You will lead multiple site teams, drive preventative maintenance, and leverage AI-driven insights to maximize uptime in a fast-paced, research-driven environment.

What You'll Do

  • Lead and manage data center site leads and their teams across multiple HPC facilities; site leads report directly to this role.
  • Recruit, mentor, and develop team members while conducting performance reviews and building a culture of operational rigor.
  • Direct onsite contractors with clear scope and validation of completed work.
  • Develop, document, and enforce operational standards for Jump's HPC data centers covering power, cooling, cabling, and hardware lifecycle.
  • Design and own the preventative maintenance program, including scheduled inspections, component replacements, and firmware/capacity reviews to minimize unplanned downtime.
  • Drive continuous improvement of operational processes and pursue automation—including AI-driven approaches—to reduce manual effort and human error.
  • Serve as the subject matter authority on HPC data center power distribution, power striping strategies, and failover/redundancy configurations.
  • Own expertise across air cooling, liquid cooling (direct-to-chip, rear-door, CDU-based), and hybrid cooling architectures.
  • Maintain deep knowledge of environmental monitoring and controls (temperature, humidity, airflow, leak detection) and ensure systems remain within design parameters.
  • Own the HPC data center monitoring strategy end-to-end: define what is monitored, set alerting thresholds, and ensure comprehensive visibility into facility and hardware health.
  • Leverage AI tools to analyze telemetry data, identify failure patterns, predict potential issues, and accelerate root cause analysis during incidents.
  • Lead critical incident response and drive root cause analysis and corrective actions to prevent recurrence.
  • Establish and track operational KPIs including availability, mean time to repair, and efficiency metrics.
  • Maintain deep, hands-on knowledge of server hardware architectures including multi-socket platforms, GPU/accelerator configurations, memory subsystems, NVMe/storage controllers, BMC/IPMI management, and firmware lifecycle.
  • Maintain deep, hands-on knowledge of network switch hardware including line cards, optics/transceivers, switch fabrics, and platform-specific diagnostics for Arista and Cisco platforms.
  • Evaluate new hardware platforms, drive hardware qualification and acceptance testing, and provide informed recommendations on hardware selection.
  • Own the overall hardware break-fix function across all HPC sites, ensuring rapid diagnosis and resolution for servers, GPUs, network equipment, storage, and facility infrastructure.
  • Diagnose complex hardware failures at the component level—CPUs, DIMMs, GPUs, NICs, PSUs, fans, drives, switch line cards, and optics—and direct the team to resolve efficiently.
  • Establish escalation paths, SLA targets, and reporting for hardware failures.
  • Own inventory processes and spares tracking across all HPC facilities, ensuring critical spares are stocked, tracked, and replenished to meet availability targets.
  • Maintain accurate asset records for all serialized and consumable inventory.
  • Collaborate with planning and vendor teams to ensure capacity and resilience across facilities.

What We're Looking For

  • A seasoned operations leader with experience running data center operations across multiple HPC facilities.
  • Deep expertise in HPC data center power distribution, power striping, and failover/redundancy.
  • Strong knowledge of air cooling, liquid cooling (direct-to-chip, rear-door, CDU-based), and hybrid cooling architectures.
  • Proficiency in environmental monitoring and controls; ability to keep systems within design parameters.
  • Experience driving monitoring strategy, defining metrics, and using AI/telemetry for predictive maintenance.
  • Excellent incident response and root cause analysis skills.
  • Hands-on server hardware knowledge (multi-socket CPUs, GPUs, NVMe storage, firmware lifecycle) and network hardware (Arista, Cisco).
  • Experience with hardware qualification, testing, and making informed hardware recommendations.
  • Demonstrated ability to lead hardware break-fix, diagnose component-level failures, and manage escalation and SLAs.
  • Strong inventory and spares management, asset tracking, and vendor planning.
  • Willingness to be on-site 5 days/week in Chicago or New York and travel as needed.

Nice to Have

  • Experience applying AI-driven automation to data center operations.
  • Familiarity with HPC workloads and associated tooling.
  • Experience with BMS/IPMI and advanced environmental controls.
  • Prior experience in vendor management and planning across multiple facilities.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →