SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Eventual is building infrastructure for Physical AI teams to search, index, and curate massive video corpora at petabyte scale. The company's open-source engine, Daft, processes 2+ PB/day for leading companies including Amazon, FAANG firms, Mobileye, and TogetherAI. Founded in 2022 and backed by $30M from Felicis, CRV, Y Combinator, and co-founders of Databricks and Perplexity, Eventual is solving a critical gap: today's data platforms (Databricks, Snowflake) were built for analytics, not video. Physical AI teams have video, lidar, radar, and sensor data scattered across object stores with no way to find what they need without weeks of human annotation. Eventual runs vision and language models over entire video corpora, allowing researchers to ask questions like "left-arm grasp failures on deformable objects" and get curated datasets in minutes instead of weeks.
As Technical Lead, Multimodal Research, you will own the technical vision behind everything Eventual understands about video. This is a senior individual contributor role—you set research direction and make architectural decisions while staying hands-on with papers, models, and experiments. No direct reports.
Key responsibilities include: owning modeling strategy across the platform (which model families, representations, and training approaches to invest in); taking approaches from prototype into production inference at corpus scale, working with data systems and storage teams; defining evaluation standards and benchmarks that models must meet before reaching customers; owning the cost curve for understanding through architecture-level decisions on distillation, cascades, routing, and quantization to keep a 10K-hour corpus at single-digit cents per hour; and translating customer research needs into scoped technical programs with clear taxonomy, model plans, datasets, and quality instrumentation.
You'll work with a small, powerful team in the SF Mission District office, 4 days/week in-person. The role involves hands-on work across the research-engineering boundary: PyTorch prototyping alongside inference performance optimization, GPU utilization, throughput, and cost analysis. You'll prove approaches in production at petabyte scale across hundreds of thousands of hours of video.
REQUIREMENTS:
- 5+ years in applied computer vision or multimodal ML
- PhD or MS in computer science, electrical engineering, robotics, or applied mathematics with a computer vision or machine learning focus (comparable publication or production record acceptable in place of degree)
- Depth in modern vision and multimodal modeling: VLMs, VQA, embeddings, representation learning, detection, tracking, segmentation, retrieval, with judgment about what is deployable today
- Hands-on training and evaluation of models at scale on real video and sensor data
- Comfort across research and engineering boundary: PyTorch prototyping alongside inference performance, GPU utilization, throughput, and cost
- Background from perception or multimodal team at self-driving, robotics, or Physical AI company; frontier research lab; or visual-data company, ideally as senior-most person on that problem
NICE TO HAVE:
- Publications at CVPR, ICCV, ECCV, NeurIPS, ICML, or ICLR
- Built or fine-tuned VLMs or other multimodal foundation models
- Long-form video, temporal reasoning, embeddings, retrieval, or content-aware indexing at scale
- Multimodal sensor data beyond RGB (lidar, radar, depth, simulation output)
- Evaluation frameworks, labeling taxonomies, large-scale annotation programs, or inference/training optimization across GPU clusters