Talent Apply
Log in
All jobs
C

Senior DevOps Engineer

Cornerstone
Hyderabad, India Posted Sep 25, 2026
On-site

About this role

Cornerstone

Senior DevOps Engineer

Hyderabad, India

req11535

We're looking for a Senior DevOps Engineer This role is Office Based, Hyderabad Office

Senior DevOps / Cloud Platform Engineer – ML & AI Infrastructure

Job Summary

We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services.

The ideal candidate will have hands-on experience with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, GitHub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments. This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost-efficient platforms for ML services across development and production environments.

In this role you will...

Key Responsibilities

  • AWS & Kubernetes Infrastructure

  • Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.

  • Manage Amazon EKS clusters, including cluster provisioning, upgrades, scaling, networking, and troubleshooting.

  • Work with AWS SageMaker, AWS Bedrock, EKS, ECS, and related AWS services.

  • Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB.

  • Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security.

  • Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.

  • Implement best practices for security, reliability, availability, and scalability.

  • CI/CD & Azure-to-AWS Migration

  • Build and maintain CI/CD pipelines for ML and AI services.

  • Develop and manage GitHub Actions and Azure DevOps pipelines using YAML.

  • Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS.

  • Automate build, test, containerization, deployment, and release processes.

  • Establish deployment strategies across development, staging, and production environments.

  • ML Service Deployment

  • Deploy and manage ML services across AWS EKS/ECS and SageMaker.

  • Build and maintain Docker containers and Kubernetes deployments.

  • Manage environment segregation and configuration across Dev, QA, and Production.

  • Develop and maintain Kubernetes manifests and Helm charts.

  • Troubleshoot ML service deployment, networking, scaling, and runtime issues.

  • App Runner to EKS Migration

  • Lead migration of existing services from AWS App Runner to Amazon EKS.

  • Containerize applications and develop Kubernetes manifests/Helm charts.

  • Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.

  • Ensure minimal service disruption during migration and establish operational best practices on EKS.

  • Self-Hosted LLM & AI Infrastructure

  • Deploy and manage self-hosted Large Language Models and inference services.

  • Work with model serving frameworks such as vLLM.

  • Design containerized infrastructure for GPU-based model serving and inference.

  • Manage model versions, deployments, configurations, and rollback strategies.

  • Support migration of ML services from managed APIs/services to self-hosted models.

  • Work with engineering teams on API integration and inference infrastructure.

  • Databricks Administration

  • Administer Databricks workspaces, clusters, permissions, and access controls.

  • Manage cluster configuration, policies, and resource utilization.

  • Support LMI Insights and related ML/AI workloads.

  • Troubleshoot Databricks infrastructure and connectivity issues.

  • Implement appropriate security and access-control practices.

  • Elasticsearch Infrastructure

  • Design, deploy, and manage Elasticsearch clusters.

  • Perform cluster sizing, scaling, configuration, and performance optimization.

  • Manage indices, mappings, retention, and data lifecycle requirements.

  • Support Kibana configuration, dashboards, and troubleshooting.

  • Monitor Elasticsearch health, capacity, and performance.

  • Monitoring, Reliability & Auto-Scaling

  • Implement monitoring and observability for Kubernetes, AWS, and ML services.

  • Use Prometheus, Grafana, and AWS CloudWatch for monitoring and alerting.

  • Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.

  • Establish proactive alerting for infrastructure and application health.

  • Perform capacity planning and resource optimization.

  • Identify opportunities for AWS infrastructure and compute cost optimization.

  • Infrastructure as Code & Automation

  • Build and maintain infrastructure using Terraform.

  • Develop reusable Terraform modules for AWS and Kubernetes infrastructure.

  • Manage Kubernetes deployments using Helm charts.

  • Automate infrastructure provisioning, configuration, deployments, and operational tasks.

  • Maintain infrastructure documentation and deployment standards.

  • Cross-Team Collaboration

  • Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams.

  • Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.

  • Coordinate infrastructure requirements for new ML models and services.

  • Participate in production incident resolution, root-cause analysis, and continuous improvement.

  • Establish engineering standards around deployment, monitoring, security, and operational ownership.

  • You've got what it takes if you have...

  • Required Skills & Experience

  • 5+ years of experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering.

  • Strong hands-on experience with AWS.

  • Strong experience administering Amazon EKS and Kubernetes in production.

  • Hands-on experience with:

  • AWS EKS

  • AWS SageMaker

  • AWS Bedrock

  • AWS ECS

  • AWS App Runner

  • Kubernetes

  • Docker

  • NGINX / AWS ALB Ingress

  • Cloudflare / Cloudflare Tunnels

  • Strong experience with Terraform and Helm.

  • Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.

  • Strong YAML scripting and Git experience.

  • Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub.

  • Experience deploying and operating ML/AI services.

  • Experience with self-hosted LLM/model serving, preferably vLLM.

  • Experience with GPU-based workloads is highly desirable.

  • Experience with Databricks administration.

  • Experience managing Elasticsearch and Kibana.

  • Experience with Prometheus, Grafana, and CloudWatch.

  • Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery.

  • Strong understanding of cloud networking fundamentals.

  • Experience with production troubleshooting, monitoring, capacity planning, and cost optimization.

  • Strong understanding of security, IAM, secrets management, and access control.

  • Preferred / Nice-to-Have Skills

  • Experience supporting Generative AI / LLM platforms.

  • Experience with GPU infrastructure and NVIDIA/CUDA environments.

  • Experience with model lifecycle and model version management.

  • Experience migrating workloads between managed cloud services and Kubernetes.

  • Experience with AWS networking such as VPC, load balancers, security groups, and Route 53.

  • Experience with API gateways and microservice architectures.

  • Experience with Python or shell scripting for infrastructure automation.

  • Experience working with Data Engineering and ML teams in a production environment.

  • What You'll Own

  • AWS ML/AI infrastructure

  • EKS cluster administration and upgrades

  • ML service deployment and production operations

  • CI/CD automation

  • App Runner → EKS migration

  • Self-hosted LLM infrastructure and vLLM

  • Databricks platform administration

  • Elasticsearch infrastructure

  • Monitoring and auto-scaling

  • Terraform and Helm-based infrastructure automation

  • Cloud cost, reliability, and performance optimization

  • Ideal Candidate

  • The ideal candidate is a hands-on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD → AWS infrastructure → EKS deployment → monitoring → scaling → production support.

  • They should be comfortable working across both traditional DevOps infrastructure and modern AI/ML infrastructure, and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform.

  • LI-Onsite

Our Culture:

  • Spark Greatness. Shatter Boundaries. Share Success. Are you ready? Because here, right now – is where the future of work is happening. Where curious disruptors and change innovators like you are helping communities and customers enable everyone – anywhere – to learn, grow and advance. To be better tomorrow than they are today.

Who We Are:

  • At Cornerstone, we believe in AI that works in the service of people, amplifying their judgment to drive high-performing, future-ready organizations forward. Cornerstone Workforce AI™, the intelligence platform for

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →