About this role
Job title: Research Engineer, Safeguards
About the Role We are seeking ML Engineers and Research Engineers to help detect and mitigate misuse of Anthropic's AI systems. As part of the Safeguards ML team, you will build systems that identify harmful use—from policy violations to coordinated attacks—and develop defenses that keep products safe as capabilities advance. This work supports Anthropic's Responsible Scaling Policy commitments.
What You'll Do
- Develop classifiers to detect misuse and anomalous behavior at scale, including synthetic data pipelines for training classifiers and methods to automatically source representative evaluations to iterate on
- Build systems to monitor harms that span multiple exchanges, such as coordinated cyber attacks and influence operations, and develop new methods for aggregating and analyzing signals across contexts
- Evaluate and improve the safety of agentic products—developing both threat models and environments to test for agentic risks, and developing and deploying mitigations for prompt injection attacks
- Conduct research on automated red-teaming, adversarial robustness, and other research that helps test for or find misuse
What We're Looking For
- 4+ years of experience in ML engineering, research engineering, or applied research, in academia or industry
- Proficiency in Python and experience building ML systems
- Comfortable working across the research-to-deployment pipeline, from exploratory experiments to production systems
- Concern about misuse risks of AI systems, and a drive to mitigate them
- Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders
- Strong candidates may also have experience with language modeling and transformers, building classifiers, anomaly detection, or behavioral ML, adversarial machine learning or red-teaming, interpretability or probes, reinforcement learning, or high-performance, large-scale ML systems
Nice to Have
- Language modeling and transformers
- Building classifiers, anomaly detection, or behavioral ML
- Adversarial machine learning or red-teaming
- Interpretability or probes
- Reinforcement learning
- High-performance, large-scale ML systems
Compensation & Benefits
- Annual Salary: $350,000 - $500,000 USD