About this role
Job title: Machine Learning Operations Engineer II
About the Role Kensho, S&P Global’s hub for AI innovation and transformation, develops and deploys ML-based solutions across the organization. The MLOps team is the de facto ML platform team responsible for enabling ML engineers with state-of-the-art processes, tooling, and infrastructure to iterate rapidly and produce production-ready models and agents. As an MLOps Engineer, you’ll help build a mature, scalable agentic platform and work across multiple teams to improve developer experience and reduce engineering toil.
What You'll Do
- Iterate on Kensho’s ML processes to develop tools, services, and frameworks that make every stage of the ML workflow robust, auditable, and usable
- Work closely with ML engineers to understand their processes, identify pain points, and form effective solutions
- Empower engineers with stable tooling to rapidly experiment and turn research into demonstrable prototypes and mature products
- Provide resources and training for ML teams on best practices, enabling efficient productionization for high-value products and services
- Evaluate, select, and champion open source and third-party solutions, driving adoption and integration into Kensho’s platform
- Ship scalable, automated processes for model fine-tuning, reinforcement learning, and evaluation of LLMs/Agents
- Improve LLM and Agentic observability to monitor production deployments, detecting performance, decay, and drift issues
- Stay at the frontier by tracking emerging tools/frameworks, promoting best practices, and strengthening the team’s technical depth
What We're Looking For
- 2+ years of experience in ML infra, ML Ops, ML Engineering, or related
- Experience managing distributed systems with Kubernetes and understanding Kubernetes concepts and trade-offs
- Cloud Platform (AWS) understanding; experience with EKS and managed ML services like Bedrock and SageMaker
- Python proficiency (we are a Python shop)
- Familiarity with distributed computing frameworks and workflow orchestration (Ray, Airflow)
- Familiarity with software engineering best practices in an ML context
- Some basic understanding of ML concepts, LLMs, and agents
- Ability to debug distributed systems across infrastructure, networking, and application layers
- Excellent communication skills to drive adoption of new tools and best practices across multiple teams
- Curious, driven, low-ego, and eager to learn across engineering disciplines
- Technologies & Tools We Use: Python, Bash, LangGraph, PyTorch, etc.
Nice to Have
- Experience building agentic applications and contributing to open-source tools
- Familiarity with monitoring/observability tooling (e.g., Prometheus)
- Comfortable working across multiple teams with diverse product workflows
Compensation & Benefits
- Base salary range: 130,000 – 175,000 USD per year
- Eligible for annual incentive bonus and equity plans
- Kensho emphasizes autonomy, continual learning, collaborative culture, and opportunities to contribute to cutting-edge AI research