About this role
Job title: Principal Platform Engineer
About the Role Riot Games is seeking a Principal Platform Engineer on the AI Efficiency team to design, build, and evolve internal platforms, automation systems, and guardrails that scale AI services and developer workflows. The role focuses on reliability, observability, deployment safety, and enabling AI-assisted engineering tooling across Riot's AI services, internal tooling, and infrastructure.
What You'll Do
- Design, implement, and evolve internal platform capabilities that make AI Efficiency services easier to build, ship, observe, secure, and operate.
- Build and maintain self-service workflows, reusable platform abstractions, and golden paths that improve developer productivity while preserving reliability, security, and governance.
- Improve platform reliability through better monitoring, alerting, observability, deployment safety, release practices, and incident readiness.
- Define and operationalize service health indicators, SLIs, SLOs, and related reliability metrics that help teams make informed tradeoffs between reliability, velocity, and cost.
- Build automation that reduces operational toil and improves mean time to detect, respond, and recover from incidents.
- Partner with engineers throughout the software development lifecycle to embed operability, production readiness, and maintainability into system design, implementation, rollout, and ongoing support.
- Improve CI/CD systems, developer workflows, and release pipelines so shipping becomes safer, faster, and more repeatable.
- Identify platform and reliability risks across distributed systems, infrastructure, service dependencies, and operational workflows, and drive durable improvements.
- Troubleshoot AI model-serving issues across frameworks, runtimes, and hardware environments, including diagnosing configuration, compatibility, and performance issues across different GPU platforms and supporting model format conversion workflows when needed.
- Design and run resilience, recovery, and failure-mode testing to validate system behavior under stress and uncover hidden weaknesses before they impact users.
- Evaluate, integrate, and operate AI-assisted engineering tools that improve code quality, reliability, security, performance, and developer productivity across the software delivery lifecycle.
- Build and evolve automation pipelines that combine conventional CI/CD systems with agentic workflows such as automated code review, bug detection, regression analysis, test generation, remediation suggestions, and workflow verification.
- Partner with engineers to introduce safe, auditable, and measurable uses of AI agents in areas such as pull request review, operational diagnostics, UI and UX validation, accessibility checks, and production readiness checks.
- Define guardrails, approval workflows, observability, reporting, and escalation paths for AI-assisted automation to ensure these systems remain safe, trustworthy, and operationally effective.
- Establish evaluation frameworks and success metrics for AI-native development tooling, including quality lift, false positive rates, latency, cost, operational risk, and impact on engineering throughput.
- Lead or contribute to incident response and post-incident improvement work for critical internal platforms and services, with a focus on systemic fixes and long-term resilience.
- Champion platform and operational excellence through documentation, runbooks, standards, and tooling that raise the engineering bar across the broader organization.
What We're Looking For
- You’re energized by making complex systems easier to use and operate, reducing cognitive load for engineers, and building paved roads instead of one-off solutions.
- You value reliability and scalable platform design, and enjoy tackling hard infrastructure and operational problems.
- You’re excited about responsibly bringing AI-native automation patterns into real engineering workflows and enabling safe, auditable AI-assisted tooling.
- You appreciate product-minded internal platforms, self-service, and golden paths, and you can collaborate across software engineers, infrastructure teams, and cross-functional partners.