SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

Anduril - Waltham, MA, United States - In-office - posted 2026-09-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 112,000 - 149,000 / annual

Anduril Industries is a defense technology company transforming U.S. and allied military capabilities through advanced technology. The Imaging team builds and fields state-of-the-art camera and sensor systems deployed to solve real security challenges, working across the full stack from bare-metal hardware and firmware to networked services and cloud integrations. You will join as a Site Reliability Engineer on the Imaging team, serving as the frontline for keeping fielded imaging systems alive. This is not a traditional product development or cloud-SRE role. You are the person field personnel and customer-support escalations turn to when a deployed system isn't behaving. You'll be the second SRE on a small, high-trust team, working directly with the lead SRE, with real room to shape how Imaging reliability and support operate as the company scales. The work is varied and often ambiguous. You might spend a morning triaging a networking failure on a system being stood up at a remote site, an afternoon walking a field operator through sensor calibration over the phone, and the back half of a week hardening a runbook so the next person never has to solve that problem live again. When a fielded system misbehaves, the root cause could be anywhere in the stack—and you're the one who narrows it down. Key responsibilities include: owning the health and uptime of deployed imaging systems and triaging issues regardless of whether the root cause is in the network, calibration, upgrades, or sensor hardware; running point on escalations as first and second line of response for issues coming through support channels and customer-support pipelines; building and maintaining runbooks, diagnostics, and self-service tooling to reduce repeat problems; holding the boundary with engineering by owning everything short of a code fix, cleanly reproducing genuine software defects and handing them off to Mission Software Engineers; and feeding reliability signals back into the product to improve observability, upgrade safety, and failure gracefully. Travel is expected approximately 15% of the time for field support and deployment windows. Successful candidates close loops and own problems until resolved, stay calm under pressure in high-stress moments, communicate proactively without waiting to be asked for updates, are comfortable with ambiguity at stack boundaries and have a methodology for narrowing down root causes systematically, and know where the line is between frontline support and true code-level defects. REQUIREMENTS: Required: - 3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems—with real ownership after systems ship, not just standing them up - Strong Linux fundamentals, including comfort troubleshooting real networking issues (IP, routing, VPNs, connectivity in constrained or field environments) - Demonstrated ability to diagnose and resolve issues across system boundaries (networking, services, hardware interaction) without always having full visibility into every component - Comfortable owning a structured on-call rotation, including scheduled after-hours and weekend coverage - Strong written and verbal communication skills, including the ability to run a remote troubleshooting session with a non-technical operator and document what happened - Eligibility to obtain and maintain a U.S. Secret clearance Preferred: - Experience supporting fielded or deployed systems, not just development environments - Experience with fielded hardware or sensor systems (EO/IR, optical, or similar), including familiarity with sensor calibration - Scripting for diagnostics and automation (Python, Bash, or similar) - Familiarity with Nix or NixOS - Familiarity with systemd service management and observability practices on Linux - Familiarity with incident tooling (PagerDuty or equivalent)

Similar roles