SlipstreamJobsFresh Startup & VC-Backed Jobs

Manager - Model Evaluation

Zoox - Foster City, CA, United States - In-office - posted 2026-09-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Zoox is developing the first ground-up, fully autonomous vehicle fleet and supporting ecosystem. This Manager role leads a high-performing team of Data Scientists, ML Validation Engineers, and Software Engineers focused on model evaluation and validation for autonomous systems. Key Responsibilities: - Lead, mentor, and scale a team while driving roadmaps, sprint execution, resource allocation, and high-throughput model releases with safety guardrails - Foster a culture of statistical excellence, healthy skepticism, proactive risk tracking, and data-driven decision-making - Define and execute end-to-end validation strategies across offline evaluation, open/closed-loop simulation, and shadow-mode fleet benchmarking - Oversee metric development and standardization with System Safety and Autonomy teams, establishing quantitative go/no-go release criteria for Behavioral Planner and Prediction ML models - Partner with Planner, Prediction, MLOps, and Developer Efficiency teams to translate behavioral requirements into measurable validation targets and optimize evaluation pipelines - Translate complex model performance trade-offs, statistical uncertainty, and safety risks into clear, data-driven recommendations for release decisions and executive leadership Requirements: - Master's or PhD in Computer Science, Robotics, Applied Statistics, or related field - 3+ years of direct engineering management experience leading Data Science, Machine Learning, or V&V engineering teams - 7+ years of technical experience in robotics, autonomous systems, or AI/ML - Strong background in ML model validation, behavioral evaluation frameworks, system-level performance benchmarking, and statistics - Strong technical foundation in Python and modern data/ML platforms - Exposure to or conceptual literacy in large-scale production codebases (C++ or distributed systems) - Proven ability to partner with systems software engineers, review technical architecture, and understand compute/performance trade-offs - Track record of leading teams evaluating complex robotic systems - Proven familiarity with modern C++/Python ML environments, simulation frameworks, and high-throughput ML evaluation pipelines - Demonstrated ability to navigate complex organizational trade-offs between release velocity, compute cost, and safety rigor Bonus Qualifications: - Experience with autonomous vehicles, robotics, or other safety-critical systems - Experience building large-scale simulation, model evaluation, or validation infrastructure - Experience with reinforcement learning, generative AI, or distributed ML systems

Similar roles