SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 184,000 - 272,000 / annual
Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI, powering commercial autonomous trucks and robotaxis. We're seeking a Senior ML Ops Engineer to build and evolve our training infrastructure and developer platforms.
In this role, you will:
• Build and evolve training infrastructure on Kubernetes, managing GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and workflow engines for reliable long-running training.
• Shape developer-facing surfaces—CLIs, SDKs, job submission, templates, and paved paths—designed collaboratively with the teams who use them. Make the common case one command while keeping the uncommon case possible.
• Shorten the inner loop: reduce time to first training run, edit-to-signal latency, and local iteration cycles. Measure and drive down these metrics systematically.
• Evangelize best-in-class tooling and frameworks. Evaluate ecosystem innovations honestly, prototype working solutions, and communicate switching costs clearly.
• Strengthen the data and artifact layer: dataset versioning, sharding, and high-throughput loading of large multimodal sensor data to saturate GPUs instead of waiting on I/O.
• Convert one-off Python scripts into durable, tested, documented, observable libraries, CLIs, and services with sane defaults.
• Make experiments legible: establish experiment hygiene, build dashboards researchers trust, implement a real model registry, and track lineage from dataset to checkpoint to simulation result.
• Ship CI/CD for models alongside autonomy and simulation, validating model changes the same way code changes are validated.
• Build observability across the ML stack: utilization, throughput, failure modes, queue times, and cost per experiment. Enable researchers to self-diagnose failures.
• Treat documentation, onboarding, and support as product surfaces. Create golden-path guides, enable new researchers to be productive on day two, and convert repeat support questions into shipped fixes.
• Drive adoption through prototyping with real users, observing workflows, and iterating. A tool nobody adopts didn't ship.
• Make the platform reliably boring: fewer failures, faster recovery, minimal manual steps.
• Build guardrails with Security, IT, and Infrastructure: access controls, data handling, and cost governance that work in IP-sensitive environments while remaining self-serve.
Qualifications:
• 5+ years of software or infrastructure engineering, including tools/platforms for other engineers and operating ML or data-intensive production systems.
• Hands-on Kubernetes expertise: GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and ability to debug clusters under load.
• Excellent Python with a track record of designing APIs and CLIs others enjoy using.
• Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar).
• Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling from a user-enablement perspective.
• Fluency with containers, CI/CD, modern build systems, and large monorepos.
• Ability to influence without authority: evaluate frameworks on merit, pilot credibly, persuade skeptical senior engineers to adopt changes.
• Collaborative default: co-own systems rather than draw boundaries around your part.
• User empathy: prioritize fixing the third-most-interesting problem blocking ten people over the most interesting one blocking nobody.
• Strong product instincts, strong writing, comfort operating autonomously in ambiguous territory.
• Passion for self-driving technologies, frontier AI, and what small world-class teams can accomplish with right infrastructure.
Bonus/Nice-to-Have:
• Internal developer platform, research platform, or DevEx work with demonstrated adoption growth.
• Large-scale distributed GPU training (hundreds to thousands of accelerators), NCCL, high-performance cluster networking, collective communication tuning.
• High-throughput loading of LiDAR or camera data; formats like Parquet or WebDataset.
• Workflow and scheduling systems: Argo Workflows, Ray, Flyte, Kubeflow, or Slurm.
• Build-system depth (Bazel or similar) with remote caching in monorepos.
• Simulation infrastructure or large-scale batch evaluation pipelines.
• Background in ML, robotics, or autonomous systems infrastructure.
• Experience in security- and IP-sensitive production environments.
• Open-source contributions to ML infrastructure or developer tools.