About this role
Moveworks is seeking a Senior Software Engineer for the Agent Eval Platform in Mountain View, California. You will own the central problem of how to score multi-step agent actions precisely enough to teach and improve behavior. The role focuses on building the judgement layer of our agent evaluation platform, including rubrics, judges, calibration against human labels, and the methodology that makes a score meaningful. The aim is to produce a calibrated signal that can both explain failures and serve as a reward signal to prevent them. This is applied ML work at the intersection of evaluating and training LLM-driven agents operating in stateful, multi-tenant enterprise environments.
Responsibilities
- Lead the evaluation orchestration at scale: design and run multi-turn agent scenarios end-to-end, including environment setup, user simulators, the user↔agent↔world loop, data collection (transcripts, traces, final state), validators, scoring, and teardown.
- Manage scheduling, retries, high-concurrency execution, and production-scale run isolation; maintain versioned specs, datasets, and reports with run-to-run comparisons as a core capability.
- Consolidate existing one-off evals into a single orchestration service with one source of truth and centralized scheduling and retry logic.
- Establish reliability floor and service-level objectives for the evaluation harness and move toward self-serve capabilities so teams can run evals without bespoke integrations.
- Build agent observability and tracing: migrate to OpenTelemetry-native observability, define a span data model for agent trajectories, ensure prompts and completions survive the pipeline, and implement fault attribution and cross-run diffing.
- Provide a robust debug surface and harness support, define tracing contracts with the agent team, and ensure traceability across async boundaries and long-lived sessions.
- Develop the stateful simulation environment: stateful fakes of enterprise systems (ITSM, HR, knowledge bases, inventory) backed by a persistent datastore; support per-run data injection and hermetic setup/teardown; implement LLM-driven user simulators and deterministic state-machine simulators; maintain contract-testing mocks against real API schemas to preserve simulation fidelity; build isolated sandbox environments that reproduce config, identity, content, and permissions for each run.
- Across all areas, lay the foundation for using evaluation signals to optimize agents, not just measure them.
Qualifications
- Experience in at least 3 of the following areas:
- Distributed systems with emphasis on determinism, reproducibility, idempotency, and isolation
- Orchestration and workflow runtimes (DAGs, scheduling, retries, backfills, high-concurrency systems)
- Observability internals (OpenTelemetry SDKs/collectors, semantic conventions, high-cardinality trace data)
- Concurrent and asynchronous programming (Python asyncio, Go concurrency, structured cancellation)
- Data-intensive pipelines (high-volume ingest, schema evolution, sampling/retention trade-offs)
- gRPC/protobuf service and interface design
- 5+ years building production backend or infrastructure systems
- Strong programming skills in Python or Go (ideally both)
- Experience designing and operating systems that handle real traffic at scale
- Ability to make non-deterministic systems measurable; you may not need an ML background, but you should be interested in turning fuzzy agent behavior into actionable signals for release gating
- Comfort with ambiguity; these are novel problems without textbook solutions