SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

Anduril - Waltham, MA, United States - In-office - posted 2026-09-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 112,000 - 149,000 / annual

Anduril Industries is a defense technology company transforming U.S. and allied military capabilities through advanced technology. The Imaging team builds and fields state-of-the-art camera and sensor systems deployed to solve real security challenges, working across the full stack from bare-metal hardware and firmware to networked services and cloud integrations. You will join as a Site Reliability Engineer on the Imaging team, serving as the frontline for keeping fielded imaging systems alive. This is not a traditional cloud-SRE role, nor a product development role. You will be the second SRE on a small, high-trust team, working directly with the lead SRE, with real room to shape how Imaging reliability and support operate as the team scales. The work is varied and often ambiguous. You might spend a morning triaging a networking failure on a system being stood up at a remote site, an afternoon walking a field operator through sensor calibration over the phone, and the back half of a week hardening a runbook so the next person never has to solve that problem live again. When a fielded system misbehaves, the root cause could be anywhere in the stack—and you are the one who narrows it down. Key responsibilities include: - Own fielded system reliability: you are responsible for the health and uptime of deployed imaging systems, triaging and diagnosing issues whether the root cause is in the network, calibration, upgrades, or sensor hardware itself. - Run point on escalations: serve as first and second line of response for issues coming through support channels and Anduril's customer-support pipeline, acting as the deep-expertise backstop. - Turn fires into runbooks: build and maintain runbooks, diagnostics, and self-service tooling that reduce repeat problems and shrink the support load over time. - Hold the boundary with engineering: own everything short of a code fix. When an issue is a genuine software defect, cleanly reproduce it, document it, and hand it off to Mission Software Engineers. - Feed reliability back into the product: turn what you learn in the field into signals that make systems more supportable—better observability, safer upgrades, more graceful failure. - Travel expected approximately 15% of the time for field support and deployment windows. You thrive in this role if you close loops (owning problems until resolved), stay calm under pressure, communicate proactively, are comfortable with ambiguity at the stack boundary, and know where the line is between frontline support and engineering-level defects. REQUIREMENTS: Required: - 3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems—with real ownership after systems ship, not just standing them up - Strong Linux fundamentals, including comfort troubleshooting real networking issues (IP, routing, VPNs, connectivity in constrained or field environments) - Demonstrated ability to diagnose and resolve issues across system boundaries (networking, services, hardware interaction) without always having full visibility into every component - Comfortable owning a structured on-call rotation, including scheduled after-hours and weekend coverage - Strong written and verbal communication skills, including the ability to run a remote troubleshooting session with a non-technical operator and document what happened - Eligibility to obtain and maintain a U.S. Secret clearance Preferred: - Experience supporting fielded or deployed systems, not just development environments - Experience with fielded hardware or sensor systems (EO/IR, optical, or similar), including familiarity with sensor calibration - Scripting for diagnostics and automation (Python, Bash, or similar) - Familiarity with Nix or NixOS - Familiarity with systemd service management and observability practices on Linux - Familiarity with incident tooling (PagerDuty or equivalent)

Similar roles