About this role
About the Role
Join a small, technical research and engineering team building rigorous benchmarks for evaluating AI agents on realistic, domain-specific workflows. You will own benchmark design and implementation, helping ensure evaluation results are reliable and useful to research and industry teams.
What You'll Do
-
Design, implement, and maintain benchmarks for evaluating AI agents on domain-specific tasks.
-
Collaborate with subject-matter experts to turn real workflows into realistic tasks and evaluation criteria.
-
Build reliable infrastructure to run models and agents against evaluation tasks at scale.
-
Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes.
-
Validate how benchmark results relate to real-world performance and evaluation needs.
-
Write clear technical documentation and reports for research and engineering audiences.
What We're Looking For
-
Two to four years of experience in software engineering, machine learning engineering, or research, including at least two years focused on AI benchmarks, evaluations, or agent environments.
-
Hands-on experience designing, implementing, and operating benchmarks or evaluation infrastructure for AI agents or large language models.
-
Proficiency with Python, Docker, and Linux environments.
-
Experience working with subject-matter experts to model workflows across technical or business domains and define evaluation criteria.
-
Experience developing metrics, statistical analyses, or validation studies for benchmark quality and real-world relevance.
-
Strong technical writing, attention to detail, and ability to work independently in an early-stage environment.
-
Experience with reinforcement learning pipelines, data generation, or agent evaluation is useful. Published work or technical writing on AI evaluation is also valued.
Compensation & Benefits
Annual salary range: $100,000 to $170,000 USD. Visa sponsorship is available.
Location
On-site in Singapore, Singapore.