SlipstreamJobsFresh Startup & VC-Backed Jobs

Technical Lead - Autonomy Evaluation

Atoms - San Francisco, CA, United States - In-office - posted 2026-09-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 185,000 - 242,000 / annual

Atoms is building Physical AI—real-world robots for industries including food, mining, and transport. The company integrates hardware, software, AI, operations, and manufacturing to deploy autonomous systems at scale. As Technical Lead for Autonomy Evaluation, you will own the metrics, platforms, and processes that determine whether autonomy releases are safe and ready for deployment. This is a hands-on technical leadership role with direct accountability for evaluation infrastructure used across the organization. Key responsibilities: • Metrics ownership: Define and own safety and behavior metrics (collision, near-miss, trajectory agreement, comfort) and the go/no-go release criteria built on them. • Log replay and scoring pipeline: Own the system that converts recorded logs into scored test cases, covering open-loop and closed-loop evaluation of perception, localization, and planning against ground truth. • Regression and benchmarking: Build and maintain the system that benchmarks new software against baselines across the log corpus for every candidate release, including dataset design, regression detection, and triage workflows. • Test coverage and data mining: Own strategies for mining the corpus for rare and important events, including sampling strategy, log selection by expected value, and coverage of underrepresented scenario classes. • Execution at scale: Ensure deterministic and reproducible evaluation execution, run traceability, version control, and cost tracking per replay-hour. • Engineering practices: Set standards for metric definitions, test design, code review, and evaluation result documentation. • Team leadership: Mentor engineers joining the function, set technical bar, and participate in hiring. You will work on-road vehicles and build infrastructure that allows each new software release, sensor configuration, and platform to be assessed seamlessly. Requirements: • Bachelor's degree in Computer Science, Robotics, Electrical Engineering, Statistics, or related field with 8+ years of relevant experience. • 2+ years as a technical lead or engineering manager for an evaluation, validation, or test team. • Ownership of an offline or system-level evaluation system for an autonomous vehicle or robotics program, or comparable large-scale ML model evaluation system in production, with accountability for metrics and release decisions. • Experience designing safety or behavior metrics for autonomous systems and defending them to engineering and executive audiences. • Hands-on experience building log replay at scale, including scoring against ground truth in open-loop and closed-loop evaluation. • Experience running evaluation pipelines on GPU clusters or comparable large-scale compute on recurring cadence, with responsibility for reproducibility, throughput, and cost. • Experience with evaluation dataset design, benchmarking, and regression detection across software or model versions. • Working command of applied statistics for system validation: statistical uncertainty, sampling, rare-event analysis, and data volume requirements. • End-to-end understanding of autonomous systems: sensing, perception, localization, planning, and controls. • Experience on an autonomous vehicle or robotics platform. • Working experience with robotics middleware and its logging and replay tooling. • Experience handling autonomous driving or robotics sensor data across multiple modalities (camera, LiDAR, radar) and timestamps. • Proficiency in Python and ability to read, navigate, and debug C++ codebases.

Similar roles