Talent Apply
Log in
All jobs
WF

Senior Site Reliability Engineer

Wikimedia Foundation
Remote (Worldwide)
Remote

About this role

About the Role The Wikimedia Foundation is looking for a Senior Site Reliability Engineer to support and develop the platform behind Wikipedia, serving millions of users worldwide. This remote-first role is part of a globally distributed SRE team that embraces open source and open documentation. Occasional travel (1-2 times per year) for in-person events and team meetings is expected.

What You'll Do

  • Perform day-to-day operational/DevOps tasks on Wikimedia’s public facing infrastructure (deployment, maintenance, configuration, troubleshooting)
  • Implement and utilize configuration management and deployment tools (Puppet, Kubernetes)
  • Lead continuous improvement by automating the installation, configuration and maintenance of services on our platform
  • Collaborate with product teams helping them bring scalable functionality to our users by assisting in the architectural design of new services and making them operate at scale
  • Participate in a 24/7 on-call rotation shared across the broader SRE team, including incident response and post-incident follow-up
  • Collaborate with a global, cross-functional team in an asynchronous communication environment
  • Mentor peers in your areas of technical and operational strength

What We're Looking For

  • 6+ years experience in an SRE/Operations/DevOps role as part of a team
  • Experience with shell and scripting languages used in an SRE context (Python, Go, Bash, Ruby; we primarily use Python) and configuration management tools (Puppet, Ansible; we use Puppet)
  • Experience with distributed caching systems and understanding of underlying algorithms and performance optimization
  • Experience with package management on Linux systems (Debian)
  • Strong Linux system-level troubleshooting skills
  • History of automating tasks and processes, identifying gaps, and finding automation opportunities
  • Strong English language skills (verbal and written) and ability to work independently in a globally distributed team across multiple time zones
  • Experience leading and participating in incident response and post-incident rituals, with root cause analysis and preventive measures

Nice to Have

  • Linux kernel tuning
  • Experience with monitoring, metrics and logging infrastructure (Prometheus, Grafana, etc.)
  • Developing/contributing to Free and Open Source software, or being part of an open-source community
  • Experience with LAMP stack technologies (PHP/HHVM, memcached/Redis) — MediaWiki experience is a definite plus
  • Experience defining cross-team SLOs and their implementation
  • Operating an on-premise filesystem or object store at scale (OpenStack Swift or Ceph)
  • Experience with other advanced distributed storage and database systems (Cassandra, MariaDB, etc.)

Compensation & Benefits

  • Salary information is not disclosed in the posting.
  • Remote-first organization with a globally distributed team, diverse workforce, and equal opportunity employment. Travel 1-2 times per year for in-person events and team meetings may be required.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →