SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Aalyria is a leading aerospace technology company supplying laser communications and temporospatial software-defined networking platforms. The company, which acquired technology from Google, specializes in satellite and airborne mesh networks, cislunar and deep-space communications, and orchestration of planetary mesh networks across land, sea, air, and space.
This is a strategic, high-impact Site Reliability Engineer role focused on building the observability and reliability infrastructure for mission-critical satellite and space systems. This is not a traditional "keep the lights on" SRE position, but rather a platform-building opportunity to create the nervous system for networks that interconnect and orchestrate satellite megaconstellations and deep-space missions.
Key responsibilities include designing and building Aalyria's centralized observability platform, integrating and scaling tools for metrics (Prometheus), logging (Loki), and distributed tracing (Tempo/OpenTelemetry). You will define and implement Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for core products to ensure launch readiness. You'll partner with software engineers to implement observability best practices, develop standard templates and documentation, and configure tooling like OpenTelemetry libraries.
Additionally, you will automate deployment, scaling, and management of the observability stack using Infrastructure as Code (Terraform) and GitOps principles (ArgoCD). You'll work closely with the core infrastructure team to ensure deep visibility into Kubernetes clusters and underlying GCP and AWS environments. You'll develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems.
The role includes on-call responsibilities and is ideal for an SRE who thrives on platform-building challenges and wants to build a production-grade observability stack from the ground up for mission-critical aerospace systems.