About this role
About the Role The Wikimedia Foundation is looking for a Senior Site Reliability Engineer to support and develop the platform behind Wikipedia, serving millions of users worldwide. This remote-first role is part of a globally distributed SRE team that embraces open source and open documentation. Occasional travel (1-2 times per year) for in-person events and team meetings is expected.
What You'll Do
- Perform day-to-day operational/DevOps tasks on Wikimedia’s public facing infrastructure (deployment, maintenance, configuration, troubleshooting)
- Implement and utilize configuration management and deployment tools (Puppet, Kubernetes)
- Lead continuous improvement by automating the installation, configuration and maintenance of services on our platform
- Collaborate with product teams helping them bring scalable functionality to our users by assisting in the architectural design of new services and making them operate at scale
- Participate in a 24/7 on-call rotation shared across the broader SRE team, including incident response and post-incident follow-up
- Collaborate with a global, cross-functional team in an asynchronous communication environment
- Mentor peers in your areas of technical and operational strength
What We're Looking For
- 6+ years experience in an SRE/Operations/DevOps role as part of a team
- Experience with shell and scripting languages used in an SRE context (Python, Go, Bash, Ruby; we primarily use Python) and configuration management tools (Puppet, Ansible; we use Puppet)
- Experience with distributed caching systems and understanding of underlying algorithms and performance optimization
- Experience with package management on Linux systems (Debian)
- Strong Linux system-level troubleshooting skills
- History of automating tasks and processes, identifying gaps, and finding automation opportunities
- Strong English language skills (verbal and written) and ability to work independently in a globally distributed team across multiple time zones
- Experience leading and participating in incident response and post-incident rituals, with root cause analysis and preventive measures
Nice to Have
- Linux kernel tuning
- Experience with monitoring, metrics and logging infrastructure (Prometheus, Grafana, etc.)
- Developing/contributing to Free and Open Source software, or being part of an open-source community
- Experience with LAMP stack technologies (PHP/HHVM, memcached/Redis) — MediaWiki experience is a definite plus
- Experience defining cross-team SLOs and their implementation
- Operating an on-premise filesystem or object store at scale (OpenStack Swift or Ceph)
- Experience with other advanced distributed storage and database systems (Cassandra, MariaDB, etc.)
Compensation & Benefits
- Salary information is not disclosed in the posting.
- Remote-first organization with a globally distributed team, diverse workforce, and equal opportunity employment. Travel 1-2 times per year for in-person events and team meetings may be required.