SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 166,000 - 220,000 / annual
Anduril Industries is seeking a Senior Site Reliability Engineer to join the Observability team, responsible for building and operating the company's production telemetry systems across ground and cloud nodes. The role involves architecting and designing a core observability plane for high-volume, high-availability telemetry ingestion, managing central monitoring and alerting infrastructure, and supporting both cloud-connected and offline/airgapped environments.
Key responsibilities include:
- Building and operating a robust, high-availability observability plane serving 100s of environments with strict uptime requirements
- Architecting systems that gracefully scale to meet demand while maintaining reliability for dependent engineering teams
- Partnering closely with internal customers (engineering teams, Robotics Data Foundation, Fleet Management) to understand needs and proactively improve systems
- Enabling metrics-driven engineering and seamless incident investigation across production systems
- Determining monitoring and alerting posture for core observability infrastructure
- Interfacing with SRE and software teams across the business to ensure seamless infrastructure integration
- Supporting agentic workflows for root-cause analysis and cross-service investigation during production incidents
The ideal candidate brings 8+ years of SRE or related production systems experience, strong familiarity with container orchestration (Docker, Kubernetes) and cloud infrastructure (AWS, GCP, Azure), and hands-on experience with observability stack components (Prometheus, Grafana, ClickHouse, Victoria Metrics, ELK). Experience in on-call rotations for high-availability systems and a user-focused engineering mindset are essential. U.S. Person status is required due to access to export-controlled data.