SlipstreamJobsFresh Startup & VC-Backed Jobs

Director, Site Reliability Engineering

Anduril - Costa Mesa, CA, United States - In-office - posted 2026-08-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Anduril Industries is a defense technology company transforming U.S. and allied military capabilities through advanced autonomy, AI, computer vision, and sensor fusion. The company's Lattice OS powers a family of systems that turn thousands of data streams into real-time 3D command and control centers. The Director of Site Reliability Engineering leads the SRE organization within CorpTech Platform, the internal engineering force multiplier behind Anduril's corporate systems. This role owns the reliability strategy for software powering Anduril's business and manufacturing operations, partnering with engineering leaders to make availability, resilience, performance, and production readiness shared outcomes across the portfolio. Key responsibilities include building and leading the SRE organization through hiring, coaching, performance management, and development of managers and senior technical leaders. You will own the reliability strategy for CorpTech Platform's production portfolio, translating operational risk and business criticality into a sequenced roadmap of engineering and infrastructure investments. The role establishes shared accountability between SRE and software engineering through clear service ownership, SLOs, production-readiness standards, on-call expectations, and escalation mechanisms. You will set direction for observability, deployment safety, incident management, capacity planning, resilience, and disaster recovery across business-critical systems and shared infrastructure. This includes partnering with engineering, product, security, and business operations leaders to make informed trade-offs between delivery speed, reliability, cost, and platform health. The role leads organizational response to critical incidents, ensures effective communication under pressure, and holds teams accountable for durable post-incident improvements. Additional focus areas include reducing recurring incidents and operational toil by turning failure patterns into platform capabilities and automation, defining meaningful reliability measures and operating reviews that expose systemic risk, and creating a healthy on-call system with appropriate staffing, tooling, and training. You will also establish reliability practices for AI-enabled systems, including observability, evaluation, degradation, and fallback mechanisms for non-deterministic production behavior.

Similar roles