About this role
About the Role
Join a small, senior team building autonomous AI agents for executive productivity. You will develop new agent capabilities and create reliable ways to evaluate them, helping ensure that important real-world tasks are completed accurately before new capabilities reach users.
What You'll Do
-
Build agents that complete delegated work end to end without requiring the user to check each step.
-
Develop trustworthy evaluations that measure how accurately agents perform different tasks.
-
Ship new agent capabilities in the tools people already use, supported by evidence that they are ready.
-
Investigate production failures and deliver durable, validated fixes.
-
Build large-scale failure detection to identify known patterns and unexpected issues.
-
Improve agent recovery when a tool or website fails.
-
Create repeatable workflows so engineers can measure whether agent changes improve performance.
What We're Looking For
-
Approximately 2 to 11 years of relevant experience, including a few years in one role with demonstrated progression.
-
Experience building or evaluating production AI or machine learning systems, especially agents or LLM-powered products.
-
Strong Python skills and familiarity with Django, React, and TypeScript.
-
Experience with evaluation and monitoring tools such as Braintrust or Raindrop is relevant.
-
Comfort working across implementation, evaluation, and iteration in a fast-moving environment.
Compensation & Benefits
Salary range: $200,000 to $300,000 USD annually. Visa sponsorship is not available.
Location
Hybrid, with four days per week in the office in Palo Alto, California, United States.