About this role
Job title: Staff DevOps Engineer
About the Role
As a Staff DevOps Engineer at Tazapay, you will design and operate highly scalable, secure cloud infrastructure for a large fintech platform. You will lead DevOps practices, mentor engineers, and drive platform initiatives across reliability, performance, and security, including PCI-DSS and SOC2 compliance. You will collaborate with product and engineering teams to translate requirements into scalable infrastructure and a clear platform roadmap.
What You'll Do
-
Make architectural tradeoffs, design and deliver highly scalable, reliable, secure and fault-tolerant cloud infrastructure
-
Be a role model for DevOps and platform engineers; mentor engineers
-
Demonstrate technical leadership and create impact across teams
-
Lead AI strategy and adoption for productivity, infrastructure automation and operational efficiency
-
Drive infrastructure and deployment practices while being secure and compliant (PCI-DSS, SOC2 and others)
-
Participate in infrastructure, architecture and security design reviews to maintain our high engineering standards
-
Partner with engineering and product management teams to define and execute the platform roadmap
-
Translate business and engineering requirements into scalable and extensible infrastructure design
-
Proactively manage stakeholder communication related to deliverables, risks, changes and dependencies
-
Coordinate with cross functional teams (Backend, Frontend, Data, Security, QA etc.) on planning and execution
-
Continuously improve platform reliability, developer experience, deployment velocity and operational excellence
-
Engage in capacity and demand planning, system performance analysis, tuning and cost optimization
-
Lead production outages, incident response, post-mortems and drive SRE practices across the engineering organisation
-
What We're Looking For
-
Education: Degree in Computer Science or equivalent (B.E/B.Tech or higher) with 10+ years of experience in DevOps, SRE, Platform or Cloud Engineering roles for large distributed systems in reputed organizations
-
Must have: Hands-on experience in designing, building and operating cloud infrastructure for large scale production systems on AWS
-
Must have: Deep knowledge of Linux as a production environment
-
Must have: Strong knowledge of distributed systems, networking, systems internals and asynchronous architectures
-
Must have: Expert in at least 1 of the following languages for automation and tooling: Go, Python, shell scripting
-
Must have: Extensive experience with AWS services (ECS, EC2, VPC, IAM, RDS, Kinesis, Secrets Manager, SSM, WAF and more)
-
Must have: Infrastructure as Code expertise with Terraform, CDK or CloudFormation
-
Must have: Strong experience with CI/CD tooling such as GitHub Actions, GitLab CI, Jenkins, and GitOps tools
-
Must have: Hands-on experience with observability stacks - Prometheus, Grafana, OpenTelemetry, distributed tracing and log aggregation
-
Must have: Experience with container infrastructure - Docker, container runtimes and image security
-
Must have: Experience administering and operating RDBMS/NoSQL systems at scale, such as Postgres, MongoDB and Redis
-
Must have: Ability to design and operate low latency services behind load balancers and API gateways
-
Must have: Strong understanding of system performance, scaling and reliability engineering
-
Must have: Possess excellent communication, sharp analytical abilities with proven design skills, able to think critically of the current platform in terms of growth and stability
-
Must have: Experience with microservice architecture and service-to-service communication patterns
-
Must have: Continuously refactor infrastructure and tooling to ensure high-quality design
-
Must have: Ability to plan, prioritize, estimate and execute platform releases with good degree of predictability
-
Must have: Ability to scope, review and refine user stories for technical completeness and to alleviate dependency risks
-
Must have: Passion for learning new things, solving challenging problems
-
Nice to have: Prior experience with fintech and payments
-
Nice to have: Expert level proficiency with Kubernetes in production
-
Nice to have: Prior experience operating infrastructure for stablecoins, blockchain or cryptocurrency platforms
-
Nice to have: AWS Cloud Certifications
-
Nice to have: Familiarity with security standards and compliance - PCI-DSS, SOC2, OWASP, static code analysis
-
Nice to have: Familiarity with data engineering services - Clickhouse, Dbt ETL
-
Nice to have: Experience working for a start-up
-
Compensation & Benefits
-
Salary: Not disclosed
-
Benefits: Not disclosed