About this role
Division EasyPay Everywhere
Minimum experience Associate
Company primary industry Financial Services
Job functional area Other
Job Description
Lesaka EasyPay provides accessible financial services to consumers across South Africa, including banking, lending, insurance, payments and value-added services. We use technology to make everyday financial services simpler, more accessible and more convenient for our customers.
We are looking for a Site Reliability & Production Support Engineer to help ensure the availability, performance and resilience of the systems that support our customers and our business.
The Opportunity
We are looking for a Site Reliability & Production Support Engineer to maintain and improve the availability, performance, and resilience of our production systems.
You will own production incidents from investigation through to resolution, implementing fixes yourself where possible and coordinating with other teams or external providers where needed.
You will also address root causes, prevent recurring issues, improve monitoring, and automate support and recovery tasks.
Key Responsibilities
-
Production support and incident resolution
-
Support production applications, APIs, databases, infrastructure, and integrations.
-
Own incidents from detection and triage through investigation, service restoration, and closure.
-
Prioritise incidents based on severity, business impact, and affected services.
-
Investigate issues using logs, metrics, traces, database queries, application code, and infrastructure diagnostics.
-
Implement configuration, script, infrastructure, and application code changes within your area of responsibility.
-
Carry out controlled rollbacks, failovers, restarts, and message reprocessing with appropriate safeguards.
-
Coordinate engineering teams and external providers where their expertise or access is needed, and drive issues through to resolution.
-
Keep incident records and communicate impact, progress, and recovery status to stakeholders.
-
Participate in an agreed on-call rotation, including support for critical incidents outside business hours.
-
Root cause analysis and prevention
-
Conduct root cause investigations and blameless incident reviews.
-
Identify causes and contributing factors across applications, infrastructure, processes, and dependencies.
-
Document incident timelines, recovery actions, findings, and preventative measures.
-
Assign owners to corrective actions, track completion, and verify that fixes address the underlying problem.
-
Identify recurring incidents and support requests, and implement changes to prevent them.
-
Reliability and observability
-
Define and monitor service level indicators and objectives with engineering and business stakeholders.
-
Build and maintain dashboards, alerts, centralised logging, and distributed tracing.
-
Monitor system health, performance, dependencies, and customer impact.
-
Reduce alert noise and improve failure detection.
-
Address reliability risks, performance bottlenecks, capacity constraints, and single points of failure.
-
Define and support resilience tests, disaster recovery exercises, and backup and restore testing.
-
Automation and operational improvement
-
Automate repetitive support tasks, diagnostics, health checks, and recovery procedures.
-
Maintain operational tools, scripts, runbooks, and troubleshooting guides.
-
Contribute to infrastructure as code, CI/CD pipelines, and deployment safeguards.
-
Ensure services have the monitoring, health checks, rollback plans, and documentation needed for production support.
-
Work with developers to improve timeout handling, retries, idempotency, and graceful failure.
-
Follow production access controls, change management, and audit requirements.
-
Required Experience and Skills
-
A Bachelor's degree or diploma in Computer Science, Information Technology, Software Engineering, Engineering, or a related technical field.
-
5+ years' relevant experience in SRE, production engineering, DevOps, technical application support, or a similar production-focused technical role.
-
Experience investigating and resolving production incidents.
-
Ability to troubleshoot applications, databases, operating systems, networks, and cloud infrastructure.
-
Experience supporting cloud-hosted and on-premises workloads.
-
Practical SQL skills, including diagnosing connectivity, query performance, locking, and connection pooling issues.
-
Experience with monitoring, logging, alerting, and distributed tracing tools.
-
Proficiency in at least one scripting or programming language, such as Python, Bash, PowerShell, C#, or Go.
-
Ability to read application code, investigate defects, and work with developers on fixes.
-
Familiarity with CI/CD, version control, infrastructure as code, and safe production deployments.
-
Clear written and verbal communication, especially during incidents.
-
Ability to take ownership, solve problems systematically, and prioritise under pressure.
-
Practical production experience is important to us. While 5+ years' relevant experience is preferred, we will consider the depth and relevance of hands-on experience alongside formal qualifications and years of experience.
-
Advantageous Experience
-
Microsoft Azure.
-
One or more of .NET, Laravel, Go, Python, or Rust.
-
Container platforms and orchestration tools, such as Docker and Kubernetes.
-
OpenTelemetry, Prometheus, Grafana, or similar tools.
-
Terraform and Git-based delivery pipelines.
-
High-volume transactional systems, distributed architectures, or customer-facing financial services platforms.
-
Disaster recovery, high availability, and multi-region architectures.
-
What Success Looks Like
-
Faster incident detection and service restoration, with fewer recurring incidents.
-
Clear ownership of incidents and corrective actions through to completion.
-
Permanent fixes that are verified to address root causes.
-
Measurable reliability objectives that inform engineering priorities.
-
Useful dashboards, actionable alerts, and up-to-date runbooks.
-
Less manual support work through automation and safe, automated recovery.
-
Timely, accurate stakeholder communication during incidents.