SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anduril Industries is seeking a Staff DevOps Engineer to architect and own the reliability infrastructure for CorpTech Platform, the internal systems powering Anduril's business and manufacturing operations. This is a high-impact, platform-wide role where you set the standards for how production systems are observed, deployed, and operated across the organization.
You will design and operate the observability platform (metrics, distributed tracing, structured logging, alerting, dashboarding) that gives engineering teams real-time production visibility at scale. You own deployment infrastructure and release-safety mechanisms, including progressive rollouts, canary analysis, automated rollback, and deployment gates that enable fast shipping without sacrificing stability.
Key responsibilities include defining SLO frameworks that make reliability measurable and actionable; identifying systemic reliability risks and driving infrastructure investments that eliminate entire failure classes; establishing production-readiness standards that embed reliability into the development lifecycle; and leading incident response for complex, multi-system failures with a focus on durable systemic improvements.
You will also develop reliability patterns for AI-enabled systems, including monitoring for model drift and graceful fallback under novel failure modes. You provide technical direction to SRE and infrastructure engineers, facilitate cross-team architecture decisions, and represent reliability concerns in platform-level planning.
Required: 10+ years in site reliability engineering, production engineering, or infrastructure engineering at architecture or platform-wide scope. Demonstrated experience designing and owning reliability infrastructure (observability platforms, deployment systems, incident management tooling) at meaningful scale. Deep technical fluency in distributed systems, Kubernetes, cloud platforms (AWS, GCP, Azure), networking, storage, and their failure modes. Proficiency in systems programming languages (Go, Python, Rust, or equivalent).