About this role
Principal Software Engineer
Microsoft is seeking a Principal Software Engineer to lead the evolution of the offline evaluation platform that gates Copilot’s quality. You will shape a vision for an AI-first, agentic-first evaluation platform where autonomous agents are not only measured but can operate, extend, and heal the system. You will tackle the hardest problem of delivering trustworthy results across deep dependencies—Copilot’s model, retrieval, and scoring services—driven far outside their normal operating envelope and under heavy and growing load. You will define end-to-end evaluation of agentic Copilot experiences, from multi-turn trajectories and tool selection to task outcomes and modalities such as UX and voice, and establish how the product defines and measures quality. This role offers the opportunity to influence the evaluation strategy for one of the world’s most visible AI products and to work at the frontier of large-scale agentic systems and reliability engineering while growing as a technical leader across a large engineering organization.
Responsibilities
- Set the technical direction for the Copilot offline evaluation platform and collaborate with teams across Copilot to translate ambiguous quality questions into rigorous, reproducible evaluations: the scenarios, metrics, and scorecards that gate what ships.
- Lead the architecture of an AI-first, agentic-first evaluation platform where autonomous agents are first-class operators, designing services, pipelines, and tooling that agents can run, extend, and reason about, with observability and guardrails that ensure trustworthy agent-driven operation.
- Own the platform’s hardest challenge: delivering trustworthy results across a deep chain of dependencies driven far outside their normal operating envelope, designing for graceful degradation, intelligent retry, dependency-aware gating, and reliability under sustained, growing load.
- Define end-to-end evaluation of Copilot’s agentic experiences, from multi-turn trajectories and tool selection to task outcomes and modalities such as UX and voice, building simulation, data collection, and scoring capabilities that make agent behavior measurable.
- Lead by example and mentor engineers across teams to build extensible, maintainable systems, driving modernization of the evaluation runtime and platform to scale with demand while improving cost and latency.
- Stay at the frontier of agentic systems, large-scale AI evaluation, and autonomous operations, bringing in new trends and patterns and sharing knowledge to raise the bar for the entire Copilot organization’s quality measurement.
Requirements
- Required: Bachelor’s Degree in Computer Science or a related technical field and 6+ years of technical engineering experience with coding in languages such as .NET, Java, JavaScript, Rust, or Python, or equivalent experience.
- Required: Ability to meet Microsoft, customer, and/or government security screening requirements, including the Microsoft Cloud Background Check and periodic rechecks.
- Preferred: Master’s Degree in Computer Science or a related technical field and 8+ years of technical engineering experience with coding in the listed languages, or Bachelor’s Degree with 12+ years of experience, or equivalent experience.
- Preferred: Deep experience with large-scale distributed systems and reliability engineering, delivering dependable outcomes across many interacting services under heavy or unusual load.
- Preferred: Experience building AI-first or agentic systems, designing systems, APIs, and tooling intended to be operated by autonomous agents, with observability, guardrails, and safety mechanisms for trustworthy operation.
- Preferred: Experience with AI/ML evaluation, measurement, or benchmarking systems, such as LLM-based grading, metric design, or large-scale data and quality pipelines.