SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer, Cyber

Anduril - Arlington, VA, United States - Hybrid - posted 2026-09-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 146,000 - 220,000 / annual

Anduril Industries is a defense technology company transforming U.S. and allied military capabilities through advanced technology. The Cyber business line is a new, fast-growing division focused on offensive cyber mission capabilities, leveraging Anduril's fleet of autonomous vehicles, Lattice OS, mesh networks, and hardware products to deploy cyber capabilities in unconventional or difficult-to-reach environments. As a Site Reliability Engineer in Anduril Cyber, you will solve a wide variety of problems involving networking, systems integration, and distributed systems while making pragmatic engineering tradeoffs. Your efforts will ensure that Anduril's software is reliable, scalable, and deployable to achieve critical national security outcomes. You will work closely with software developers, customers, and external vendors to deliver working offensive cyber products. You will own the full deployment pipeline—from CI/CD and pre-production test environments, through canary deployments in customer-hosted integration environments, to production in air-gapped enclaves. You will also serve as a steward of Anduril's mission and technological advantage in customer meetings, explaining system behavior and absorbing requirements firsthand. Site Reliability Engineers must be driven by a "Whatever It Takes" mindset—executing in an expedient, scalable, and pragmatic way while keeping the mission top-of-mind. Key responsibilities include: owning the health of deployed systems and keeping them running with minimal downtime; automating and improving software deployment processes into air-gapped, TS/SCI environments; designing, building, and maintaining CI/CD and automated test infrastructure for complex hardware and software systems; developing metrics dashboards, TUIs, scripts, and tools that automate deployment steps or help debug the software stack; driving engineering requirements based on onsite observations; performing root cause analysis and diagnosing issues in mission-critical systems across the software stack, Lattice OS stack, and external vendor services; building strong relationships with internal and external customers to identify technical solutions; and driving continuous improvement by instrumenting systems, analyzing failures, and leading post-mortem events spanning software, firmware, and hardware. REQUIREMENTS: - Currently possesses and is able to maintain an active U.S. TS/SCI security clearance - Based in the DC metro area to support 3-5 days per week working on site at customer facilities - 4+ years of experience in a Sys Admin, Site Reliability, DevOps, or Software Engineering role - Deep, practical experience with Linux and Kubernetes (or a similar container orchestrator) - Working knowledge of network fundamentals and the ability to debug connectivity in a locked-down environment - Experience delivering and maintaining systems on air-gapped and security-hardened networks - Strong proficiency in Python or Bash for automation and debugging, and the ability to read and debug service code in a compiled language such as Go - Excellent written and verbal communication skills for collaborating with a cross-functional engineering team and external customers PREFERRED QUALIFICATIONS: - Experience in debugging and resolving networking issues - Ability to quickly understand and navigate complex, multi-disciplinary systems and established codebases - Experience building automation for hardware-in-the-loop (HIL) or software-in-the-loop (SIL) test environments - Ability to drive consensus across internal and external stakeholders

Similar roles