SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior / Staff ML Ops Engineer

Waabi - Remote - Remote - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 184,000 - 272,000 / annual

Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI, powering commercial autonomous trucks and robotaxis. We're seeking a Senior ML Ops Engineer to build and evolve our training infrastructure and developer platforms. In this role, you will: • Build and evolve training infrastructure on Kubernetes, managing GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and workflow engines for reliable long-running training. • Shape developer-facing surfaces—CLIs, SDKs, job submission, templates, and paved paths—designed collaboratively with the teams who use them. Make the common case one command while keeping the uncommon case possible. • Shorten the inner loop: reduce time to first training run, edit-to-signal latency, and local iteration cycles. Measure and drive down these metrics systematically. • Evangelize best-in-class tooling and frameworks. Evaluate ecosystem innovations honestly, prototype working solutions, and communicate switching costs clearly. • Strengthen the data and artifact layer: dataset versioning, sharding, and high-throughput loading of large multimodal sensor data to saturate GPUs instead of waiting on I/O. • Convert one-off Python scripts into durable, tested, documented, observable libraries, CLIs, and services with sane defaults. • Make experiments legible: establish experiment hygiene, build dashboards researchers trust, implement a real model registry, and track lineage from dataset to checkpoint to simulation result. • Ship CI/CD for models alongside autonomy and simulation, validating model changes the same way code changes are validated. • Build observability across the ML stack: utilization, throughput, failure modes, queue times, and cost per experiment. Enable researchers to self-diagnose failures. • Treat documentation, onboarding, and support as product surfaces. Create golden-path guides, enable new researchers to be productive on day two, and convert repeat support questions into shipped fixes. • Drive adoption through prototyping with real users, observing workflows, and iterating. A tool nobody adopts didn't ship. • Make the platform reliably boring: fewer failures, faster recovery, minimal manual steps. • Build guardrails with Security, IT, and Infrastructure: access controls, data handling, and cost governance that work in IP-sensitive environments while remaining self-serve. Qualifications: • 5+ years of software or infrastructure engineering, including tools/platforms for other engineers and operating ML or data-intensive production systems. • Hands-on Kubernetes expertise: GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and ability to debug clusters under load. • Excellent Python with a track record of designing APIs and CLIs others enjoy using. • Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar). • Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling from a user-enablement perspective. • Fluency with containers, CI/CD, modern build systems, and large monorepos. • Ability to influence without authority: evaluate frameworks on merit, pilot credibly, persuade skeptical senior engineers to adopt changes. • Collaborative default: co-own systems rather than draw boundaries around your part. • User empathy: prioritize fixing the third-most-interesting problem blocking ten people over the most interesting one blocking nobody. • Strong product instincts, strong writing, comfort operating autonomously in ambiguous territory. • Passion for self-driving technologies, frontier AI, and what small world-class teams can accomplish with right infrastructure. Bonus/Nice-to-Have: • Internal developer platform, research platform, or DevEx work with demonstrated adoption growth. • Large-scale distributed GPU training (hundreds to thousands of accelerators), NCCL, high-performance cluster networking, collective communication tuning. • High-throughput loading of LiDAR or camera data; formats like Parquet or WebDataset. • Workflow and scheduling systems: Argo Workflows, Ray, Flyte, Kubeflow, or Slurm. • Build-system depth (Bazel or similar) with remote caching in monorepos. • Simulation infrastructure or large-scale batch evaluation pipelines. • Background in ML, robotics, or autonomous systems infrastructure. • Experience in security- and IP-sensitive production environments. • Open-source contributions to ML infrastructure or developer tools.

Similar roles