SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anduril Industries is seeking a Site Reliability Engineer to join the Imaging team, which builds and deploys state-of-the-art camera and sensor systems for U.S. and allied military applications. This is a specialized SRE role focused on keeping fielded systems alive in production environments, not traditional cloud infrastructure work.
You will be the second SRE on a small, high-trust team working directly with the lead SRE. Your primary responsibility is ensuring the health and uptime of deployed imaging systems across the full stack—from bare-metal hardware and firmware to networked services and cloud integrations. When fielded systems experience issues, you are the frontline diagnostician, triaging problems that could originate anywhere in the stack: networking failures, sensor calibration issues, software defects, or hardware interactions.
Key responsibilities include: owning fielded system reliability and uptime; serving as the escalation point for support tickets and customer issues; converting recurring problems into durable runbooks and self-service diagnostics; maintaining the boundary between field support and product engineering by cleanly reproducing and documenting software defects; and feeding field reliability insights back into product design to improve observability, upgrade safety, and failure handling.
The work is varied and often ambiguous. You might triage a networking failure at a remote deployment site in the morning, walk a field operator through sensor calibration over the phone in the afternoon, and spend the rest of the week hardening runbooks to prevent future incidents. You'll own an on-call rotation including after-hours and weekend coverage. Approximately 15% travel is expected for field support and deployment windows.
Required qualifications: 3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems with real post-deployment ownership; strong Linux fundamentals and networking troubleshooting skills (IP, routing, VPNs, constrained environments); demonstrated ability to diagnose issues across system boundaries without full visibility; comfort with structured on-call rotations; strong written and verbal communication for remote troubleshooting and documentation; eligibility to obtain and maintain a U.S. Secret clearance.