Talent Apply
Log in
All jobs
A

Senior Software Engineer, Interpretability Infrastructure

Anthropic
San Francisco, CA
HybridUSD 315,000 - 560,000 / year

About this role

About the Role Join Anthropic's Interpretability team to build and maintain the specialized infrastructure for researching and auditing large language models. You will design and optimize the tooling that enables researchers to instrument model internals, run forward/backward passes, and apply interventions at scale to support safety and interpretability work.

What You'll Do

  • Build and maintain the specialized inference and training infrastructure powering interpretability research, including instrumented passes, activation extraction, and steering vector application.
  • Profile, optimize, and resolve scaling and efficiency bottlenecks in distributed systems and hardware utilization.
  • Design tools, abstractions, and platforms enabling researchers to experiment rapidly without engineering blockers.
  • Help bring interpretability research into production safety audits with real deadlines and high reliability expectations.
  • Collaborate across the stack from model internals and accelerator-level optimization to user-facing research tooling.
  • Translate research needs into engineering solutions in partnership with researchers.

What We're Looking For

  • 5-10+ years of professional software development experience.
  • Proficiency in at least one major language (e.g., Python, Rust, Go, Java) and strong Python fluency.
  • Highly curious, able to learn unfamiliar domains quickly and apply new knowledge to engineering work.
  • Ability to prioritize high-impact work, tolerate ambiguity, and challenge assumptions.
  • Preference for fast-moving, collaborative projects over solo efforts.
  • Genuine interest in interpretability and AI safety; research experience not required.
  • Commitment to ethical and societal implications of your work.
  • Comfort working closely with researchers to translate research needs into engineering solutions.

Nice to Have

  • Experience optimizing large-scale distributed systems.
  • Language modeling fundamentals with transformers.
  • High-performance LLM optimization: memory management, compute efficiency, parallelism, inference throughput.
  • Hands-on work with PyTorch/CUDA on GPUs or JAX/XLA on TPUs.
  • Track record of building tooling to support research teams or performing research with engineering challenges.
  • Representative projects such as instrumenting LLMs for activations, streaming data pipelines of activations, and steering-inference systems.

Compensation & Benefits

  • Annual salary: $315,000 - $560,000 USD.
  • Location-based hybrid work policy: SF office presence required at least 25% of the time; exceptional candidates may be considered for remote arrangements.
  • Other benefits available per company policy.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →