Talent Apply
Log in
All jobs
H

Founding ML Engineer

HiringCafe
Cupertino, CA Posted Jul 16, 2026
On-siteUSD 250,000 - 310,000 / year

About this role

About the Role Founding ML Engineer to turn powerful AI/ML models into fast, reliable production systems. You will own the bridge between model development and user-facing infrastructure: deploying models, optimizing latency and throughput, and scaling serving systems. This is a hands-on role focused on production performance, GPU utilization, inference architectures, and reliability. What You'll Do

  • Deploy and integrate researcher-trained model checkpoints into our cloud infrastructure and production pipelines.
  • Profile and benchmark model performance to identify latency, throughput, memory, and compute bottlenecks.
  • Implement optimization techniques such as quantization, pruning, batching, caching, efficient attention, and precision trade-offs while preserving model quality.
  • Build scalable multi-GPU inference systems for search, ranking, recommendations, agents, and other AI-powered product experiences.
  • Design reliable model-serving architecture that can support millions of users.
  • Develop efficient training and fine-tuning workflows where needed, including distributed training, mixed precision, and parallelism strategies.
  • Work closely with our search & engineering teams to make model deployment a smooth part of our development workflow. What We're Looking For
  • Have deployed and optimized deep learning models in production environments.
  • Have experience with large-scale model serving, multi-GPU inference, or high-throughput inference systems.
  • Understand inference optimization techniques such as quantization, pruning, compilation, batching, caching, and memory optimization.
  • Have strong instincts for profiling, benchmarking, and debugging model performance.
  • Are familiar with efficient attention mechanisms, transformer optimization, or modern LLM/embedding/ranking model infrastructure.
  • Have worked with inference frameworks or serving stacks such as SGLang, vLLM, TensorRT, or equivalent.
  • Can write clean, production-quality code and integrate ML systems into backend infrastructure.
  • Are comfortable with cloud platforms, distributed systems, storage systems, and modern ML training or serving workflows.
  • Want ownership, leverage, and responsibility from day one. Compensation & Benefits
  • Salary: $250k-$310k per year + equity
  • We offer generous health, dental, and vision coverage, paid parental leave, and relocation support.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →