About this role
About Eventual
Robots and world models learn about the physical world from video: millions of hours of people cooking, building, carrying and fixing things. That footage is piling up faster than anyone can review it, and much of it is mislabeled, out of sync or not useful for training. Most teams still choose what to train on by having people watch a small sample. When a model learns from bad examples, it learns the wrong behavior.
Eventual is building data curation for Physical AI. We process video and sensor data at petabyte scale and extract signals from it: how hands and objects interact, what happens over time (a failed grasp, a retry, a recovery) and whether a clip is reliable enough to learn from. Leading Physical AI labs and robotics teams use these signals to decide what goes into their next training run.
The work spans large-scale data systems, computer vision and ML research, and it ships directly into how frontier models get trained. We’re a small team from AWS, Lyft and Tesla, backed by $30M from Felicis, CRV, Y Combinator and the co-founders of Databricks and Perplexity. We helped power the last generation of Physical AI in self-driving, and now we’re building what the next generation trains on.
Join our small (but powerful!) team, 4 days/week in our SF Mission District office.
Our Mission:
Our goal is to build Scenario Mining and Data Curation for robot fleet data. We empower Physical AI and robotics teams to instantly find, curate, and stream the data they need to train frontier models.
Eventual is an agile team where every engineer has high ownership across the stack from our compute infrastructure, to our data storage/querying layers and model training/deployment.
Your Role:
As a Software Engineer on the Data Systems team, you will build key capabilities for Eventual. We build storage and a high-throughput data engine over petabytes of video, lidar and high-frequency telemetry data to power frontier robotics and Physical AI labs. You will work directly on the architecture powering real-time indexing of perception/robotics data, distributed storage/compute, and dataloading at line-rate to GPUs for model training and inference. We operate as a tight-knit, experienced engineering team that values technical autonomy, deep execution, and a passion for solving hard distributed systems problems.
Key Responsibilities:
-
Multimodal Storage (Data Lake): built against modern columnar data lake formats (Apache Parquet, Apache Iceberg etc) optimized for high-dimensional video, lidar and sensor logs.
-
Query Engine: build powerful querying capabilities. Our multi-stage query systems are built on database fundamentals such as partitioning, indexing, query planning, embeddings/vector search for search and retrieval as well as LLMs/VLMs for perception-based query predicates.
-
Dataloading: Improve memory stability, throughput, and zero-copy data flow through streaming computation and line-rate CUDA tensor delivery to GPUs.
What we look for:
-
Proven track record building resilient, high-throughput distributed systems or database engines using Rust or C++.
-
3+ years of experience diving deep into engine internals—such as vectorized execution, query planning/optimization, distributed task scheduling, or zero-copy networking.
-
Practical exposure to scaling cloud infrastructure (AWS S3) and managing heavy-compute data pipelines (bonus points for experience with CUDA, GPU streaming, or video decoding frameworks).
-
High agency and adaptability to thrive in an autonomous, fast-paced startup environment building cutting-edge infrastructure for frontier robotics
Perks & Benefits
-
In-person tight knit team with 4x a week in office
-
Competitive comp and startup equity
-
Catered lunches and dinners for SF employees
-
Commuter benefit
-
Team building events & poker nights
-
Health, vision, and dental coverage
-
Flexible PTO
-
Latest Apple equipment
-
401k plan with match!