SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Eventual is a Series A AI infrastructure company ($30M raised from Felicis, CRV, Microsoft M12, and others) building Daft, a distributed data engine purpose-built for multimodal AI. The company is solving a critical bottleneck: today's data platforms (Databricks, Snowflake) were designed for analytics, not the petabyte-scale video, lidar, and sensor corpora that power Physical AI systems like humanoid robots and autonomous vehicles. Robotics and video-AI teams currently lose 20-40% of training time to data loading alone, while GPU bandwidth grows 2-3× per generation but storage and pipelines lag behind.
Daft is already in production at Amazon (2 PB/day), major FAANG companies (60-100 PB), Mobileye, TogetherAI, and CloudKitchens. The company is now building a video-native index and streaming layer that delivers curated datasets to GPUs at line rate, saturating B200s today and targeting NVL72 and Vera Rubin tomorrow.
As a Systems Engineer on the Dataloading team, you will design and build the video-native dataloader that transforms multi-petabyte corpora into GPU-ready tensors at line rate. You'll profile and optimize the full data path from object store through NVMe, page cache, host RAM, to device RAM—eliminating every avoidable copy and stall. You'll work with the top Physical AI labs training on the newest generation hardware (H100, B200, GB200, NVL72), ensuring GPUs stay saturated on real customer training jobs while enabling the complex sampling patterns researchers need without sacrificing model FLOPs utilization.
Key responsibilities include designing rank-aware, NVMe-cached dataloaders with random access into clips; profiling and optimizing the full data path; saturating latest hardware on customer training jobs; owning performance benchmarks against baselines (custom DataLoaders, DALI, decord, LeRobot); partnering with researchers at partner labs to integrate the loader into their training stacks; and collaborating cross-team on index/format boundaries and model-output ingestion.
You should have obsessive attention to systems-level performance, strong opinions on io_uring, deep expertise in Rust/C++/C, and solid OS fundamentals (page cache, scheduling, syscalls, NUMA, memory hierarchies). You need a strong intuition for where bytes travel across NVMe, memory, network, PCIe, and NVLink, and their throughput/latency budgets. GPU experience is a plus but not required on day one; the company will upskill you on CUDA, SLURM, and NVL72. Nice-to-haves include SLURM/Kubernetes experience, CUDA expertise, deep memory/caching knowledge, video decode pipeline experience, and open-source systems contributions.
The team is small but world-class, assembled from AWS, Render, Pinecone, and Tesla. You'll work 4 days/week in the SF Mission office with competitive compensation, meaningful equity, catered meals, commuter benefits, health/vision/dental coverage, flexible PTO, and latest Apple equipment.