About this role
CONSULTANT
Publication Date: Oct 6, 2026
Ref. No: 551527 Location: Chennai, IN skillset : Java , Observability (ELF , Grafan , Splunk) , Github , Service now(ticketing) , AWS/GCP Replacement of Harikaran.
Roles & Responsibilities:
- Provide hands-on support for the runtime operation of our applications, ensuring high availability and performance.
- Collaborate with software engineering and infrastructure teams to troubleshoot and resolve runtime issues, including performance bottlenecks, scalability challenges, and system failures.
- Contribute to the design and implementation of monitoring, alerting, and logging solutions to proactively identify and address potential runtime issues.
- Participate in incident response and root cause analysis efforts to ensure the stability and resilience of the applications.
- Work closely with cross-functional teams to understand application requirements and provide input on runtime and operational considerations during the software development lifecycle.
- Contribute to the development and maintenance of runtime automation and tooling to streamline operational processes and improve efficiency.
- Develop common framework components (to be leveraged by enterprise applications), define standards for configuration, monitoring, reliability, and performance engineering
- Create automation and ensure automated tests are completed for new features.
- Good attitude, communication, willingness to learn and collaborate.
- Continuously improve automated remediation tasks to ensure the highest levels of availability.
- Cloud: Manage secure, scalable, and highly available cloud infrastructure.
- Kubernetes & Containers: Deploy, operate, and troubleshoot containerized workloads.
- Observability: Implement monitoring, logging, tracing, dashboards, and actionable alerts.
- Reliability: Define SLOs/SLIs, manage error budgets, and improve service availability.
- Networking: Troubleshoot DNS, TCP/IP, HTTP/S, TLS, routing, and load balancing.
- Linux & Systems: Administer and troubleshoot Linux systems and performance issues.
- Programming/Scripting: Automate operational tasks using Python, Go, or Shell.
- Infrastructure as Code: Provision and manage infrastructure using Terraform or equivalent IaC tools.
- CI/CD: Build and maintain automated, reliable deployment pipelines.
- Incident Management: Respond to incidents, perform RCA, and implement preventive actions.