Talent Apply
Log in
All jobs
JT

HPC Operations Engineer

Jump Trading
Sydney, Australia; Mumbai, India
On-site

About this role

Job title: HPC Operations Engineer

About the Role Jump Trading's global High Performance Computing Team is looking to add an HPC Operations Engineer in our Sydney and Mumbai offices. The role focuses on maintaining and improving Linux HPC environments at scale to support data pipelines and quantitative research, requiring hands-on, adaptable approach and a proactive mindset.

What You'll Do

  • Provide front-line operational support for 24/7 Linux HPC compute, storage, and interconnects.
  • Work with RDMA fabrics, parallel filesystems, HPC batch schedulers, FUSE filesystems, internal Jump software, multi-vendor hardware, cybersecurity requirements, and demanding client workloads.
  • Solve problem reports and questions from Jump's research community, escalating as needed and managing the entire problem lifecycle.
  • Respond to alerts in a timely fashion.
  • Participate in large, coordinated maintenance operations, including evenings and weekends.
  • Work on global projects across a wide range of infrastructure.
  • Write code for diagnosing, resolving, triaging difficult problems, and automating frequently performed tasks.
  • Collaborate across teams to write code and testing infrastructures in multiple languages.
  • Manage relationships with outside vendors, including domestic and international travel to meet current and potential vendors.
  • Implement and support performance monitoring and fault monitoring systems.
  • Develop and improve systems and user documentation.
  • Develop and monitor tools used to maintain a production computing environment.
  • Provide operational support as your primary job function.
  • Adhere to company cybersecurity and IT policies, including using approved hardware and software.
  • Participate in an on-call rotation.
  • Other tasks as assigned or needed.
  • Work from company office an average of 5 days a week.
  • Be willing to work a maintenance window on Friday evening or Saturday morning on a rotating basis.

What We're Looking For

  • A desire for operational work as the primary job function.
  • At least 2+ years of professional experience with Linux systems.
  • HPC experience including parallel filesystems (Lustre, GPFS), batch systems (Slurm, Grid Engine), and high-performance network interconnects is a plus, but not required.
  • High proficiency with at least one programming/scripting language (e.g., Go, Python, C) and the ability to learn additional languages quickly.
  • Ability to perform root cause analysis.
  • Strong verbal and written communication skills, including the ability to communicate effectively and efficiently with both coworkers and third-party vendors.
  • Strong collaboration skills with a willingness to undertake tasks of various technologies and complexities.
  • Ability to independently manage complex projects and multiple workstreams.
  • Strong sense of urgency.
  • Willingness to perform regular operational maintenance work during evenings and weekends and as needed.
  • Ability to work effectively in a busy, open floor plan office environment.
  • Reliable and predictable availability.

Nice to Have

  • HPC experience including Lustre/GPFS, and other parallel filesystems, plus experience with HPC interconnects and batch systems is a plus.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →