Talent Apply
Log in
All jobs
M

Senior Software Engineer

Microsoft
United States, Washington
USD 119,800 - 261,000 / year

About this role

A Senior Software Engineer at Microsoft designs, builds, and operates large-scale distributed systems that analyze fleet telemetry, identify security and freshness risks, and drive automated remediation across Microsoft 365 infrastructure. The role sits at the intersection of cloud engineering, operating systems, firmware, security, data platforms, and AI-assisted operations to improve the trustworthiness and reliability of services used by millions of customers daily. You will partner closely with security teams, hardware and firmware engineering, Azure infrastructure and Fleet management teams, service owners, and site reliability engineers to deliver solutions that improve platform health, reduce operational risk, and keep Microsoft 365 infrastructure secure, compliant, and up to date. This position is based in Redmond, Washington and requires in-office presence a minimum of three days per week.

Microsoft’s mission is to empower every person and every organization to achieve more. We value growth, collaboration, respect, integrity, and accountability, and we strive to create an inclusive culture where everyone can thrive at work and beyond.

Responsibilities

  • Design and build large-scale distributed services that improve operating system, firmware, security, and patch freshness across Microsoft 365 infrastructure.
  • Develop telemetry, analytics, and reporting platforms to measure fleet compliance, update readiness, vulnerability exposure, firmware health, and remediation progress.
  • Build intelligent services that detect freshness gaps, identify at-risk infrastructure, prioritize remediation, and forecast fleet security or compliance risk.
  • Create safe automation workflows for operating system updates, firmware upgrades, security patch deployment, configuration enforcement, and remediation orchestration.
  • Partner with security, hardware, firmware, operating system, and infrastructure teams to define and execute data-driven strategies for improving fleet health and reducing exposure.
  • Build AI-assisted experiences that speed investigation of patching failures, firmware issues, configuration drift, vulnerability exposure, and live-site incidents.
  • Improve operational safety through staged rollout systems, health gates, rollback mechanisms, policy enforcement, and strong observability.
  • Participate in architecture reviews, code reviews, design discussions, and live-site operations.
  • Mentor engineers and contribute to engineering excellence across the organization.
  • Drive projects from concept and design through production deployment, measurement, and operational ownership.

Qualifications

  • Required: Bachelor’s Degree in Computer Science or a related technical field and 4+ years of technical engineering experience with coding in languages such as C, C++, C#, Java, JavaScript, or Python, or equivalent experience. Ability to meet Microsoft, customer, and/or government security screening requirements, including the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
  • Preferred: Master’s Degree in Computer Science or a related technical field and 6+ years of experience with coding in the listed languages, or Bachelor’s Degree with 8+ years of experience, or equivalent experience.
  • In-depth knowledge of operating system internals, with strong expertise in Windows and practical Linux experience, including debugging and troubleshooting kernels, drivers, system services, networking stacks, crash dumps, and performance issues.
  • Experience collecting, correlating, and analyzing diagnostic data at scale (event logs, traces, performance counters, networking telemetry, driver diagnostics, and other system-level signals) to identify root causes and drive remediation. Ability to build or use data pipelines (e.g., Kusto/Azure Data Explorer) to detect systemic issues and drive remediation across the M365 Fleet.
  • Experience diagnosing complex OS issues using tools such as WinDbg, Windows Performance Analyzer, ETW, and analyzing crash dumps, driver failures, networking, and performance.
  • Tier 2 or Tier 3 United States Government clearance to work in secure Microsoft cloud environments.
  • Experience with security patching, vulnerability management, compliance reporting, configuration management, or secure infrastructure operations.
  • Experience developing automation systems for safe deployment, remediation, rollout orchestration, rollback, or fleet-wide policy enforcement.
  • Experience applying data science, statistics, forecasting, anomaly detection, or machine learning to infrastructure health, security posture, or operational risk problems.
  • Experience developing AI-powered operational tools, investigation assistants, or agent-based solutions.
  • Demonstrated ability to independently drive complex technical projects from concept through production deployment.
  • Strong collaboration and communication skills with the ability to influence across teams and organizations.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →