SlipstreamJobsFresh Startup & VC-Backed Jobs

ML Infra Engineer, Modeling

Physical Intelligence - San Francisco, CA, United States - In-office - posted 2026-09-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Physical Intelligence is building general-purpose AI systems for the physical world, developing foundation models and learning algorithms to power robots and physically-actuated devices. The ML Infrastructure team is seeking an experienced engineer to scale and optimize the company's training systems and core model code. In this role, you will own critical infrastructure for large-scale training, managing GPU/TPU compute, job orchestration, and building reusable, efficient JAX training pipelines. You'll work closely with researchers and model engineers to translate research ideas into experiments and production training runs. Key responsibilities include: - Designing and maintaining systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging - Scaling distributed training across TPU and GPU clusters using JAX with minimal friction - Profiling and optimizing memory usage, device utilization, throughput, and distributed synchronization - Building abstractions for launching, monitoring, debugging, and reproducing experiments - Partnering with researchers to translate infrastructure needs and guide best practices for training at scale - Contributing to core training code to support new architectures, modalities, and evaluation metrics You should have strong software engineering fundamentals and hands-on experience building ML training infrastructure or internal platforms. Preferred qualifications include large-scale training experience with JAX or PyTorch, familiarity with distributed training and multi-host setups, and experience managing training workloads on cloud platforms (SLURM, Kubernetes, GCP TPU/GKE, AWS). Bonus experience includes ML systems background (training compilers, runtime optimization, custom kernels), GPU/TPU performance tuning, robotics or multimodal models background, and experience designing abstractions that balance researcher flexibility with system reliability. This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure.

Similar roles