SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior AI Infrastructure Engineer, Physical Infrastructure

Anduril - Costa Mesa, CA, United States - In-office - posted 2026-09-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 166,000 - 220,000 / annual

Anduril Industries is seeking a Senior AI Infrastructure Engineer to lead the design, deployment, and operation of large-scale GPU compute infrastructure supporting the company's AI and autonomy initiatives. This is a hands-on role within the CorpTech Infrastructure Engineering team, which provides foundational infrastructure enabling engineers, researchers, and product teams across Anduril to deploy fast, scalable systems. You will own the full lifecycle of GPU cluster robustness and resilience. Key responsibilities include: physically racking, stacking, and cabling GPU compute systems (H200/B200/B300, NVL72); building and tuning high-performance interconnect fabrics (NVLink, InfiniBand, RoCE, Spectrum-X) connecting hundreds of GPUs into low-latency training and inference clusters; integrating high-performance parallel storage (VAST, DDN, Weka) to sustain throughput for distributed training and multi-modal datasets; automating cluster deployment and configuration end-to-end using infrastructure-as-code; operating and extending Kubernetes/Run:AI environments for GPU scheduling, quota management, and multi-tenant workload isolation; owning fleet health through monitoring, alerting, and rapid triage of hardware and network faults; onboarding engineers and researchers onto the platform and serving as their escalation point for infrastructure-related bottlenecks; and partnering with product-facing teams to translate emerging compute needs into platform capabilities. Required qualifications include 10+ years of hands-on infrastructure, HPC, or datacenter engineering supporting GPU compute at scale; direct experience with H200/B200/B300 GPU systems including bring-up, cabling, and firmware/driver management; experience with high-performance interconnects in clusters of hundreds of GPUs; experience with high-performance parallel storage systems; Kubernetes expertise; strong automation and infrastructure-as-code background; ability to perform physical datacenter work including lifting 50+ lbs; and eligibility to obtain and maintain a U.S. Top Secret clearance. Preferred qualifications include NVIDIA NVL72 rack-scale systems experience, LLM token serving/inference infrastructure support, network fabric tuning for RoCE/InfiniBand at scale, GPU/network observability tooling familiarity, and experience supporting infrastructure as a shared platform for multiple internal teams.

Similar roles