SlipstreamJobsFresh Startup & VC-Backed Jobs

Fielded Site Reliability Engineer

Anduril - Waltham, MA, United States - Hybrid - posted 2026-09-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 112,000 - 149,000 / annual

Anduril Industries is seeking a Site Reliability Engineer to join the Imaging team, which builds and fields state-of-the-art camera and sensor systems deployed for U.S. and allied military security. This is a fielded-systems SRE role—not traditional cloud SRE and not product development. You will be the frontline for keeping deployed imaging systems operational, serving as the escalation point when field personnel and customer support encounter issues with systems in production environments. You will be the second SRE on a small, high-trust team, working directly with the lead SRE and shaping how Imaging reliability and support scale. The work is varied and often ambiguous: you might triage a networking failure at a remote deployment site, walk a field operator through sensor calibration over the phone, or harden runbooks to prevent recurring issues. Key responsibilities include: - Own the health and uptime of deployed imaging systems, triaging and diagnosing issues across the full stack (network, calibration, upgrades, sensor hardware). - Run point on escalations from support channels and customer-support pipelines, acting as the deep-expertise backstop. - Convert recurring fires into runbooks, diagnostics, and self-service tooling to reduce repeat problems and support load. - Hold the boundary with engineering: own everything short of a code fix; cleanly reproduce and document genuine software defects for handoff to Mission Software Engineers. - Feed field reliability signals back into the product to improve observability, upgrade safety, and failure gracefully. - Travel approximately 20% of the time for field support and deployment windows. You thrive when diagnosing live systems under pressure, being the reason a deployment stays up, and turning recurring fire-drills into durable fixes. Success requires closing loops (owning problems until resolved), staying calm under pressure, communicating proactively, being comfortable with ambiguity at stack boundaries, and knowing when to hand off to engineering. REQUIREMENTS: Required: - 3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems, with real ownership after systems ship (not just standing them up). - Strong Linux fundamentals, including comfort troubleshooting real networking issues (IP, routing, VPNs, connectivity in constrained or field environments). - Demonstrated ability to diagnose and resolve issues across system boundaries (networking, services, hardware interaction) without always having full visibility into every component. - Comfortable owning a structured on-call rotation, including scheduled after-hours and weekend coverage. - Strong written and verbal communication skills, including the ability to run a remote troubleshooting session with a non-technical operator and document what happened. - Eligibility to obtain and maintain a U.S. Secret clearance. Preferred: - Experience supporting fielded or deployed systems, not just development environments. - Experience with fielded hardware or sensor systems (EO/IR, optical, or similar), including familiarity with sensor calibration. - Scripting for diagnostics and automation (Python, Bash, or similar). - Familiarity with Nix or NixOS. - Familiarity with systemd service management and observability practices on Linux. - Familiarity with incident tooling (PagerDuty or equivalent).

Similar roles