About this role
About the Role The L2/L3 Production Support Engineer will provide 24x7 production support for Java/J2EE applications in a large enterprise environment, focusing on monitoring, incident handling, RCA, and ensuring application uptime and reliability. The role requires collaboration with development, QA, infrastructure, and business teams to resolve production issues and enable smooth releases. This role also involves on-call rotations and incident bridges. What You'll Do
- Provide L2/L3 production support for Java/J2EE applications in a 24x7 enterprise environment.
- Monitor applications using Splunk, Dynatrace, AppDynamics, Kibana.
- Handle P1/P2/P3 incidents, perform deep root cause analysis (RCA), and implement preventive fixes.
- Debug issues across Java code, application servers, databases, batch jobs, and integrations.
- Support production releases, deployments, and post-release validation.
- Work with development, QA, infrastructure, and business teams to resolve production defects.
- Manage incident, problem, and change tickets using ServiceNow / Remedy.
- Analyze logs, thread dumps, heap dumps, and performance metrics.
- Ensure SLA compliance, application uptime, and system reliability.
- Participate in on-call rotations and incident bridges. What We're Looking For
- Experience providing L2/L3 production support for Java/J2EE applications in a 24x7 enterprise environment.
- Proficiency with monitoring and incident management tools: Splunk, Dynatrace, AppDynamics, Kibana.
- Strong ability to handle P1/P2/P3 incidents, perform root cause analysis (RCA), and implement preventive fixes.
- Debugging skills across Java code, application servers, databases, batch jobs, and integrations.
- Experience supporting production releases, deployments, and post-release validation.
- Excellent collaboration with development, QA, infrastructure, and business teams to resolve production defects.
- Experience managing incidents, problems, and changes using ServiceNow or Remedy.
- Ability to analyze logs, thread dumps, heap dumps, and performance metrics.
- Commitment to SLA compliance, uptime, and reliability.
- Willingness to participate in on-call rotations and incident bridges.