About this role
Cornerstone
Senior DevOps Engineer
Hyderabad, India
req11535
We're looking for a Senior DevOps Engineer This role is Office Based, Hyderabad Office
Senior DevOps / Cloud Platform Engineer – ML & AI Infrastructure
Job Summary
We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services.
The ideal candidate will have hands-on experience with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, GitHub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments. This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost-efficient platforms for ML services across development and production environments.
In this role you will...
Key Responsibilities
-
AWS & Kubernetes Infrastructure
-
Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.
-
Manage Amazon EKS clusters, including cluster provisioning, upgrades, scaling, networking, and troubleshooting.
-
Work with AWS SageMaker, AWS Bedrock, EKS, ECS, and related AWS services.
-
Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB.
-
Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security.
-
Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.
-
Implement best practices for security, reliability, availability, and scalability.
-
CI/CD & Azure-to-AWS Migration
-
Build and maintain CI/CD pipelines for ML and AI services.
-
Develop and manage GitHub Actions and Azure DevOps pipelines using YAML.
-
Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS.
-
Automate build, test, containerization, deployment, and release processes.
-
Establish deployment strategies across development, staging, and production environments.
-
ML Service Deployment
-
Deploy and manage ML services across AWS EKS/ECS and SageMaker.
-
Build and maintain Docker containers and Kubernetes deployments.
-
Manage environment segregation and configuration across Dev, QA, and Production.
-
Develop and maintain Kubernetes manifests and Helm charts.
-
Troubleshoot ML service deployment, networking, scaling, and runtime issues.
-
App Runner to EKS Migration
-
Lead migration of existing services from AWS App Runner to Amazon EKS.
-
Containerize applications and develop Kubernetes manifests/Helm charts.
-
Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.
-
Ensure minimal service disruption during migration and establish operational best practices on EKS.
-
Self-Hosted LLM & AI Infrastructure
-
Deploy and manage self-hosted Large Language Models and inference services.
-
Work with model serving frameworks such as vLLM.
-
Design containerized infrastructure for GPU-based model serving and inference.
-
Manage model versions, deployments, configurations, and rollback strategies.
-
Support migration of ML services from managed APIs/services to self-hosted models.
-
Work with engineering teams on API integration and inference infrastructure.
-
Databricks Administration
-
Administer Databricks workspaces, clusters, permissions, and access controls.
-
Manage cluster configuration, policies, and resource utilization.
-
Support LMI Insights and related ML/AI workloads.
-
Troubleshoot Databricks infrastructure and connectivity issues.
-
Implement appropriate security and access-control practices.
-
Elasticsearch Infrastructure
-
Design, deploy, and manage Elasticsearch clusters.
-
Perform cluster sizing, scaling, configuration, and performance optimization.
-
Manage indices, mappings, retention, and data lifecycle requirements.
-
Support Kibana configuration, dashboards, and troubleshooting.
-
Monitor Elasticsearch health, capacity, and performance.
-
Monitoring, Reliability & Auto-Scaling
-
Implement monitoring and observability for Kubernetes, AWS, and ML services.
-
Use Prometheus, Grafana, and AWS CloudWatch for monitoring and alerting.
-
Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.
-
Establish proactive alerting for infrastructure and application health.
-
Perform capacity planning and resource optimization.
-
Identify opportunities for AWS infrastructure and compute cost optimization.
-
Infrastructure as Code & Automation
-
Build and maintain infrastructure using Terraform.
-
Develop reusable Terraform modules for AWS and Kubernetes infrastructure.
-
Manage Kubernetes deployments using Helm charts.
-
Automate infrastructure provisioning, configuration, deployments, and operational tasks.
-
Maintain infrastructure documentation and deployment standards.
-
Cross-Team Collaboration
-
Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams.
-
Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.
-
Coordinate infrastructure requirements for new ML models and services.
-
Participate in production incident resolution, root-cause analysis, and continuous improvement.
-
Establish engineering standards around deployment, monitoring, security, and operational ownership.
-
You've got what it takes if you have...
-
Required Skills & Experience
-
5+ years of experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering.
-
Strong hands-on experience with AWS.
-
Strong experience administering Amazon EKS and Kubernetes in production.
-
Hands-on experience with:
-
AWS EKS
-
AWS SageMaker
-
AWS Bedrock
-
AWS ECS
-
AWS App Runner
-
Kubernetes
-
Docker
-
NGINX / AWS ALB Ingress
-
Cloudflare / Cloudflare Tunnels
-
Strong experience with Terraform and Helm.
-
Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.
-
Strong YAML scripting and Git experience.
-
Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub.
-
Experience deploying and operating ML/AI services.
-
Experience with self-hosted LLM/model serving, preferably vLLM.
-
Experience with GPU-based workloads is highly desirable.
-
Experience with Databricks administration.
-
Experience managing Elasticsearch and Kibana.
-
Experience with Prometheus, Grafana, and CloudWatch.
-
Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery.
-
Strong understanding of cloud networking fundamentals.
-
Experience with production troubleshooting, monitoring, capacity planning, and cost optimization.
-
Strong understanding of security, IAM, secrets management, and access control.
-
Preferred / Nice-to-Have Skills
-
Experience supporting Generative AI / LLM platforms.
-
Experience with GPU infrastructure and NVIDIA/CUDA environments.
-
Experience with model lifecycle and model version management.
-
Experience migrating workloads between managed cloud services and Kubernetes.
-
Experience with AWS networking such as VPC, load balancers, security groups, and Route 53.
-
Experience with API gateways and microservice architectures.
-
Experience with Python or shell scripting for infrastructure automation.
-
Experience working with Data Engineering and ML teams in a production environment.
-
What You'll Own
-
AWS ML/AI infrastructure
-
EKS cluster administration and upgrades
-
ML service deployment and production operations
-
CI/CD automation
-
App Runner → EKS migration
-
Self-hosted LLM infrastructure and vLLM
-
Databricks platform administration
-
Elasticsearch infrastructure
-
Monitoring and auto-scaling
-
Terraform and Helm-based infrastructure automation
-
Cloud cost, reliability, and performance optimization
-
Ideal Candidate
-
The ideal candidate is a hands-on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD → AWS infrastructure → EKS deployment → monitoring → scaling → production support.
-
They should be comfortable working across both traditional DevOps infrastructure and modern AI/ML infrastructure, and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform.
-
LI-Onsite
Our Culture:
- Spark Greatness. Shatter Boundaries. Share Success. Are you ready? Because here, right now – is where the future of work is happening. Where curious disruptors and change innovators like you are helping communities and customers enable everyone – anywhere – to learn, grow and advance. To be better tomorrow than they are today.
Who We Are:
- At Cornerstone, we believe in AI that works in the service of people, amplifying their judgment to drive high-performing, future-ready organizations forward. Cornerstone Workforce AI™, the intelligence platform for