SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Apptronik is a Series B human-centered robotics company developing AI-powered humanoid robots (Apollo) to support humanity across manufacturing, logistics, healthcare, and beyond. We're seeking a Staff MLOps Engineer to own the technical direction of our MLOps platform—the system of record connecting teleoperation data collection to deployed autonomy on Apollo robots.
In this hands-on technical leadership role, you will set the architecture for the platform layer above the training cluster, owning dataset lifecycle, experiment tracking, model registry, evaluation harnesses, and the serving/packaging path that delivers trained policies to robots in the field. You will lead by influence across MLOps, Autonomy, Data Platform, and TeleOp teams, establishing standards, contracts, and tooling that transform one-off research code into repeatable, auditable pipelines from data to deployed model.
Key responsibilities include:
**Platform Architecture & Ownership**: Define subsystem interfaces, drive architecture decisions, and establish engineering standards for how datasets, experiments, and models move through Apptronik's systems. Serve as the primary technical point of contact for cross-functional teams on model lifecycle and platform contracts.
**Dataset Lifecycle & Versioning**: Design and operate the dataset layer end-to-end—versioning, lineage, splits, and labeling integration. Ensure every trained model traces back to the exact data and code that produced it.
**Model Registry & Artifact Management**: Build and operate a first-class model registry with versioned artifacts, metadata, evaluation results, lineage, and approval workflows. Define the promotion path from "trained" to "qualified" to "deployed to robot."
**Evaluation & Qualification Harnesses**: Define offline benchmarks, simulation rollouts, and policy-gating harnesses that models must pass before reaching Apollo. Develop the metrics framework that the autonomy team trusts to gate releases.
**Serving, Packaging & Deployment**: Own the path from registered model to running inference on Apollo—packaging (ONNX, TensorRT, torch.compile), on-robot versioning, rollback, and observability of deployed policy behavior. Coordinate with Connect and Data Platform on the deploy-and-telemetry seam from the fleet.
**Mentorship & Cross-Functional Leadership**: Mentor mid-level and senior engineers on the MLOps team through code review, design review, and collaboration. Partner with the Training Infrastructure engineer on cluster/platform contracts and influence research workflows across Autonomy.
Required qualifications: Deep proficiency in Python and at least one systems-level language (Go, Rust, C++). Proven experience owning and delivering an MLOps platform end-to-end at a company shipping models to production. Expertise across the model lifecycle: dataset versioning (DVC, LakeFS, Delta), experiment tracking (MLflow, W&B, Determined), model registry, and policy serving. Strong background designing service-oriented systems on Kubernetes.