Talent Apply
Log in
All jobs
JT

HPC Production Engineer

Jump Trading
Sydney, Australia
On-site

About this role

Job title: HPC Production Engineer

About the Role Jump's global HPC Team is looking to add an HPC Production Engineer in our Sydney office. This role focuses on designing, deploying, and maintaining high performance compute and storage systems to support data pipelines and quantitative research, with a hands-on emphasis on Linux environments and scalable tooling.

What You'll Do

  • Design, implement, maintain, and support high performance compute and storage systems
  • Implement and support performance monitoring and fault monitoring systems
  • Monitor systems and storage performance, up to and including network components
  • Build tooling to compile, package, install, and upgrade software and operating system components at scale
  • Collaborate with team members and across teams to write code and testing infrastructures spanning multiple programming languages
  • Develop and improve systems and user documentation
  • Participate in large, coordinated maintenance operations, including evenings and weekends
  • Work on global projects across a wide range of infrastructure
  • Collaborate directly with researchers to optimize their use of HPC infrastructure
  • Develop and monitor the tools used to maintain a production computing environment
  • Provide operational support on a rotating basis and as needed
  • Manage relationships with outside vendors, including traveling domestically and internationally
  • Adhere to all company cybersecurity and IT policies, including performing all work using only approved hardware and software
  • Other duties as assigned or needed

What We're Looking For

  • 5+ years of professional experience in HPC, including parallel filesystems such as Lustre or GPFS, batch systems such as Slurm or Grid Engine; experience with high-performance network interconnects is a plus but not required
  • 5+ years of experience with Linux systems administration
  • High proficiency with at least one programming/scripting language such as Go, Python, or C
  • Extensive experience designing, building, and maintaining complicated, interdependent, and distributed systems
  • Extensive experience profiling and debugging application stacks (debuggers and profilers)
  • Experience with system configuration management tools such as SaltStack, Ansible, Puppet, etc.
  • A compulsion to perform root cause analysis
  • Reliable and predictable availability

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →