SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 191,000 - 253,000 / annual
Anduril Industries is a defense technology company building advanced autonomous systems and AI-powered military capabilities. The Air Dominance & Strike team develops aerial and multi-domain robotic systems including Fury (unmanned fighter jet) and Barracuda (air-breathing cruise missile), along with Lattice for Mission Autonomy—a software platform enabling masses of robots to collaborate across missions.
You will own critical components of Anduril's end-to-end machine learning platform and MLOps tooling. This role focuses on building and operating infrastructure required to train, evaluate, host, and serve complex AI models (including LLMs, computer vision, and reinforcement learning agents) across cloud environments and air-gapped, edge-deployed networks. Working closely with AI Research Scientists and Platform Engineers, you will eliminate friction in model development, optimize hardware utilization, and ensure robust delivery of models into safety-critical operational environments.
Key responsibilities include:
- Build, optimize, and maintain scalable training, orchestration, and experimentation infrastructure to accelerate state-of-the-art model development
- Identify and resolve bottlenecks in the ML lifecycle by developing tooling for experiment tracking, automated profiling, and hyperparameter tuning
- Implement and scale robust data pipelines (ETL) to process multi-modal data (video feeds, radar, flight telemetry, simulation logs) from physical assets and test sites
- Deploy high-throughput, low-latency model serving frameworks optimized for both cloud and resource-constrained, air-gapped tactical edge hardware
- Develop robust CI/CD pipelines for ML models, including automated regression testing, validation benchmarks, and safe rollout/rollback strategies
- Implement pipelines for model evaluation, validation, and reinforcement learning alignment loops (RLHF/DPO) for mission-critical deployments
- Partner with AI Researchers and Computer Vision engineers to translate modeling requirements into scalable infrastructure
- Mentor peers, conduct design and code reviews, and champion engineering best practices
REQUIREMENTS:
- 5+ years of software engineering experience with demonstrated success building and operating production-scale machine learning infrastructure or distributed systems
- Proficiency in Python, Go, or C++, with strong grasp of software engineering fundamentals, systems design, and concurrent programming
- Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks (PyTorch Distributed, Ray, Slurm, Megatron-LM)
- Experience building and maintaining distributed data pipelines handling large-scale unstructured or multi-modal datasets
- Track record of owning projects end-to-end—from technical design to production deployment and operational monitoring
- Eligible to obtain and maintain an active U.S. Top Secret security clearance
PREFERRED:
- Experience deploying ML infrastructure, model serving, or artifacts in secure, air-gapped, or regulated environments (IL5/IL6, GovCloud)
- Hands-on experience profiling GPU/accelerator workloads, resolving hardware/network bottlenecks, and optimizing compute utilization
- Experience supporting workloads for Large Language Models, Generative AI, or Reinforcement Learning pipelines
- Experience with multi-tenant cluster management, including fair scheduling, GPU slicing, and quota enforcement
- Familiarity with production ML observability frameworks, including data drift detection and automated evaluation pipelines