About this role
About Eventual
Robots and world models learn about the physical world from video: millions of hours of people cooking, building, carrying and fixing things. That footage is piling up faster than anyone can review it, and much of it is mislabeled, out of sync or not useful for training. Most teams still choose what to train on by having people watch a small sample. When a model learns from bad examples, it learns the wrong behavior.
Eventual is building data curation for Physical AI. We process video and sensor data at petabyte scale and extract signals from it: how hands and objects interact, what happens over time (a failed grasp, a retry, a recovery) and whether a clip is reliable enough to learn from. Leading Physical AI labs and robotics teams use these signals to decide what goes into their next training run.
The work spans large-scale data systems, computer vision and ML research, and it ships directly into how frontier models get trained. We’re a small team from AWS, Lyft and Tesla, backed by $30M from Felicis, CRV, Y Combinator and the co-founders of Databricks and Perplexity. We helped power the last generation of Physical AI in self-driving, and now we’re building what the next generation trains on.
Join our small (but powerful!) team, 4 days/week in our SF Mission District office.
Your Role
As Technical Lead, Multimodal Research, you’ll own the execution of our technical vision behind everything Eventual can understand about a video. Physical AI teams have video, lidar, radar, and sim outputs scattered across object stores with no way to find what they need without weeks of human annotation. Eventual runs vision/language models and pipelines over every clip in a corpus along axes the customer cares about (gripper type, failure mode, object class, scene, motion density), so a researcher can ask “left-arm grasp failures on deformable objects” and get a curated dataset in minutes.
You’ll decide which models, representations, and evaluation methods get us there, and prove them in production at petabyte scale, over hundreds of thousands of hours of video. This is a senior individual contributor role rather than a management one: you set research direction and make the architectural decisions, while staying hands-on with papers, models, and experiments.
Key Responsibilities
-
Own modeling strategy across the platform rather than one customer’s taxonomy: which model families, representations, and training approaches we invest in, which get prototyped, and when to move off one.
-
Take approaches from prototype into production inference at corpus scale, working with the data systems and storage teams on what the index and the loader require.
-
Define the evaluation standard — the benchmarks and acceptance criteria any model meets before it reaches a customer.
-
Own the cost curve for understanding: architecture-level decisions on distillation, cascades, routing, and quantization that keep a 10K-hour corpus at single-digit cents per hour of video.
-
Translate customer research needs into scoped technical programs — taxonomy, model plan, datasets, quality instrumentation — and set the technical direction for multimodal work across the company. No direct reports.
What we look for
-
5+ years in applied computer vision or multimodal ML.
-
PhD or MS in computer science, electrical engineering, robotics, or applied mathematics with a computer vision or machine learning focus. A comparable publication or production record is acceptable in place of the degree.
-
Depth in modern vision and multimodal modeling — VLMs, VQA, embeddings, representation learning, detection, tracking, segmentation, retrieval — with judgment about what is deployable today rather than competitive on a leaderboard.
-
Hands-on training and evaluation of these models at scale on real video and sensor data, and comfort across the research and engineering boundary: PyTorch prototyping alongside inference performance, GPU utilization, throughput, and cost.
-
Background from a perception or multimodal team at a self-driving, robotics, or Physical AI company, a frontier research lab, or a visual-data company, ideally as the senior-most person on that problem.
Nice to have
-
Publications at CVPR, ICCV, ECCV, NeurIPS, ICML, or ICLR.
-
Built or fine-tuned VLMs or other multimodal foundation models.
-
Long-form video, temporal reasoning, embeddings, retrieval, or content-aware indexing at scale.
-
Multimodal sensor data beyond RGB — lidar, radar, depth, or simulation output.
-
Evaluation frameworks, labeling taxonomies, or large-scale annotation programs, or inference and training optimization across large GPU clusters.
Perks & Benefits
-
In-person, tight-knit team — 4 days/week in our SF Mission office.
-
Competitive comp and meaningful startup equity.
-
Catered lunches and dinners for SF employees.
-
Commuter benefit.
-
Team-building events and poker nights.
-
Health, vision, and dental coverage.
-
Flexible PTO.
-
Latest Apple equipment.
-
401(k) plan with match.