SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

Anduril - Costa Mesa, CA, United States - In-office - posted 2026-08-11

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Anduril Industries is seeking a Staff Site Reliability Engineer to architect and own the reliability infrastructure for CorpTech Platform, the internal systems powering Anduril's business and manufacturing operations. This is not a ticket-response role—you will define how production systems are observed, deployed safely, and scaled reliably across the organization. You will set reliability architecture standards for observability (metrics, distributed tracing, structured logging, alerting, dashboarding), deployment systems (progressive rollouts, canary analysis, automated rollback), and incident management. You will design and govern SLO frameworks that create shared language between SRE, product engineering, and leadership for reliability trade-offs. Your work will eliminate entire failure classes through systemic infrastructure investments rather than patching individual symptoms. Key responsibilities include establishing production-readiness standards that embed reliability into the development lifecycle, leading incident response for complex multi-system failures, and developing reliability patterns for AI-enabled systems (monitoring for model drift, non-deterministic degradation, graceful fallback). You will drive capacity planning and cost optimization for critical paths and shared infrastructure, provide technical direction to SRE and infrastructure engineers, and represent reliability concerns in platform-level planning. The role carries broad organizational reach across CorpTech Platform's engineering teams. You will influence how software teams build and operate systems, set standards governing production readiness, and make infrastructure decisions affecting every service in the portfolio. You bring both deep systems expertise and judgment to prioritize reliability investments at each growth stage. Required: 10+ years in SRE, production engineering, or infrastructure engineering at architecture/platform scope. Demonstrated experience designing and owning reliability infrastructure (observability platforms, deployment systems, incident management tooling) at scale. Deep technical fluency in distributed systems, Kubernetes, cloud platforms (AWS/GCP/Azure), networking, storage, and failure modes. Proficiency in systems programming languages (Go, Python, Rust, or equivalent) for building production infrastructure and tooling.

Similar roles