About this role
About the Role Founding ML Engineer to turn powerful AI/ML models into fast, reliable production systems. You will own the bridge between model development and user-facing infrastructure: deploying models, optimizing latency and throughput, and scaling serving systems. This is a hands-on role focused on production performance, GPU utilization, inference architectures, and reliability. What You'll Do
- Deploy and integrate researcher-trained model checkpoints into our cloud infrastructure and production pipelines.
- Profile and benchmark model performance to identify latency, throughput, memory, and compute bottlenecks.
- Implement optimization techniques such as quantization, pruning, batching, caching, efficient attention, and precision trade-offs while preserving model quality.
- Build scalable multi-GPU inference systems for search, ranking, recommendations, agents, and other AI-powered product experiences.
- Design reliable model-serving architecture that can support millions of users.
- Develop efficient training and fine-tuning workflows where needed, including distributed training, mixed precision, and parallelism strategies.
- Work closely with our search & engineering teams to make model deployment a smooth part of our development workflow. What We're Looking For
- Have deployed and optimized deep learning models in production environments.
- Have experience with large-scale model serving, multi-GPU inference, or high-throughput inference systems.
- Understand inference optimization techniques such as quantization, pruning, compilation, batching, caching, and memory optimization.
- Have strong instincts for profiling, benchmarking, and debugging model performance.
- Are familiar with efficient attention mechanisms, transformer optimization, or modern LLM/embedding/ranking model infrastructure.
- Have worked with inference frameworks or serving stacks such as SGLang, vLLM, TensorRT, or equivalent.
- Can write clean, production-quality code and integrate ML systems into backend infrastructure.
- Are comfortable with cloud platforms, distributed systems, storage systems, and modern ML training or serving workflows.
- Want ownership, leverage, and responsibility from day one. Compensation & Benefits
- Salary: $250k-$310k per year + equity
- We offer generous health, dental, and vision coverage, paid parental leave, and relocation support.