SlipstreamJobsFresh Startup & VC-Backed Jobs

Perception Deployment Engineer - Model Deployment & Optimization

Zoox - Foster City, CA, United States - In-office - posted 2026-09-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Zoox is developing the first ground-up, fully autonomous vehicle fleet. The Perception team is pioneering multi-modality foundation models to drive next-generation autonomous system intelligence. As a Perception Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to the on-vehicle stack. You will work with experts in compressing, accelerating, and deploying complex computer vision and foundation models for power- and thermal-constrained vehicle SOCs. Key responsibilities: - Design and develop production-level, low-latency, memory-safe C++ and CUDA code for real-time perception algorithms on vehicle systems - Optimize large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, and mixed-precision inference frameworks - Architect and implement model conversion and compilation pipelines using TensorRT for edge deployment - Perform rigorous parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries - Develop and optimize custom ML OPs and TensorRT Plugins with efficient CUDA kernels to minimize latency and maximize memory bandwidth on AI accelerators Requirements: - Production-level C++ (14/17/20) and Python programming skills, with experience developing concurrent, memory-safe, real-time inference code for edge devices - Deep expertise in model compression technologies (quantization: PTQ and QAT) and mixed-precision inference frameworks (INT8, FP8, BF16/FP16) - Proven experience optimizing large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs/VLAs) utilizing Efficient Attention mechanisms (FlashAttention, Linear Attention) and KV-cache optimization (PagedAttention) - Extensive experience with model conversion/compilation pipelines (ONNX, TensorRT, torch.compile) and rigorous latency benchmarking and model quality parity validation - Proficiency in low-level programming for AI accelerators, specifically developing and optimizing custom ML OPs and TensorRT Plugins with efficient CUDA kernel implementations Bonus qualifications: - Familiarity with state-of-the-art autonomous driving perception algorithms (temporal 3D object detection, BEV, 3D Occupancy Networks) and multi-modal sensor processing (Vision, LiDAR, Radar) - Experience with end-to-end autonomous driving paradigms (VLM/VLA models, Foundation models) and edge deployment technologies (TensorRT-LLM)

Similar roles