SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 166,000 - 220,000 / annual
Anduril Industries is seeking a Senior Infrastructure Reliability Engineer to join its Infrastructure Reliability Engineering (IRE) team in Costa Mesa. The IRE team is responsible for the infrastructure and operations supporting core developer tools and on-premises compute platforms used across the entire engineering organization, including source control, CI/CD, artifact management, and on-prem infrastructure for simulation and GPU workloads.
In this role, you will own the full lifecycle of critical services—patching, upgrades, backups, scaling, and incident response—that the engineering organization depends on daily. You'll blend DevOps, SRE, and software engineering practices, with high ownership and company-wide impact. The position emphasizes continuous improvement and automation of manual, repetitive tasks.
Key responsibilities include:
- Serving as a primary owner for critical services, including on-call rotation and knowledge-sharing
- Managing the lifecycle of self-hosted developer tools (RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray)
- Designing and implementing automated systems for patching, backups with validation, and upgrades
- Scaling infrastructure to support rapid engineering organization growth
- Using Infrastructure-as-Code (Terraform) to manage environments
- Operating and troubleshooting Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
- Defining and maintaining SLOs for service availability, reliability, and performance
- Building monitoring, alerting, and observability for developer tool services
- Leading incident response and root cause analysis
- Cross-functional collaboration with platform, security, infrastructure, and software teams
Required qualifications include experience operating infrastructure outside managed cloud services (bare-metal Kubernetes, VMware ESXi/vSphere), production Docker and Kubernetes systems, strong Linux fundamentals (RHEL, Ubuntu), proficiency with at least one cloud platform, Infrastructure-as-Code tools (Terraform/OpenTofu), configuration management (Ansible, Puppet, Chef), scripting/software development (Python, Go, Bash), CI/CD pipeline familiarity, and end-to-end system ownership. Eligibility to obtain and maintain a U.S. Secret security clearance is required.
Preferred qualifications include experience with RKE2 or other bare-metal Kubernetes distributions, GitHub Enterprise Server, JFrog Artifactory, CircleCI, GitOps workflows (ArgoCD/FluxCD), highly available internal tools, security best practices and compliance, large scaling engineering organizations, monitoring platforms (Datadog, Prometheus, Grafana), SRE or hybrid SWE/DevOps backgrounds, on-prem infrastructure operations, and GPU/HPC infrastructure experience.