Talent Apply
Log in
All jobs
A

GPU Performance Engineer

Anthropic
Hybrid (office-based at least 25% of the time)
HybridUSD 280,000 - 850,000 / year

About this role

Job title: GPU Performance Engineer

About the Role Anthropic is building reliable, interpretable, and steerable AI systems. As a GPU Performance Engineer, you’ll architect and implement the foundational systems that power Claude and push the frontiers of large language models. You’ll maximize GPU utilization and inference efficiency at unprecedented scale, developing cutting-edge optimizations that enable new model capabilities and dramatically improve throughput. You’ll work at the intersection of hardware and software, spanning from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization.

What You'll Do

  • Architect and implement foundational GPU performance systems powering Claude.
  • Maximize GPU utilization and inference efficiency at unprecedented scale.
  • Develop and implement state-of-the-art kernel and algorithm optimizations.
  • Work across the stack—from low-level tensor core optimizations to coordinating thousands of GPUs in lockstep.
  • Collaborate with world-class researchers and engineers to shape AI infrastructure.
  • Deliver transformative performance improvements in production ML systems.
  • Partner with hardware vendors to influence accelerator capabilities and software stacks.

What We're Looking For

  • Deep experience with GPU programming and optimization at scale.
  • Track record of delivering measurable performance breakthroughs in production ML.
  • Ability to navigate complex systems from hardware interfaces to high-level ML frameworks.
  • Enjoy collaborative problem solving and pair programming.
  • Interest in state-of-the-art language models with real-world impact; care about societal implications; thrive in ambiguous environments.

Nice to Have

  • GPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention; tensor core optimization.
  • ML Compilers & Frameworks: PyTorch/JAX internals, torch.compile, XLA, custom operators.
  • Performance Engineering: Kernel fusion, memory bandwidth optimization, Nsight profiling.
  • Distributed Systems: NCCL, NVLink, collective communication, model parallelism.
  • Low-Precision: INT8/FP8 quantization, mixed-precision techniques.
  • Production Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration.
  • Representative projects: Co-design attention mechanisms for next-gen hardware, develop custom kernels for quantization formats, design distributed communication strategies for multi-node clusters, optimize end-to-end training and inference pipelines, build performance modeling frameworks, implement kernel fusion strategies, create resilient planet-scale distributed training infrastructure, profile and eliminate bottlenecks in production serving, partner with hardware vendors to influence future accelerators.

Compensation & Benefits

  • Annual salary: USD 280,000 – 850,000.
  • Visa sponsorship available for this role.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →