SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Engineer, Cloud Infrastructure and Networking

Skylo - Remote - Remote - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 125,000 - 135,000 / annual

Skylo has pioneered a standards-based approach to satellite connectivity, connecting smartphones and IoT devices directly to satellites with no special hardware required. The company's direct-to-device service is live on millions of activated devices across five continents, covering more than 72 million square kilometers in partnership with leading satellite operators, mobile network operators, and Tier-1 chipset makers. As Senior Engineer, Cloud Infrastructure and Networking, you will serve as the cloud infrastructure domain authority within Skylo's production Non-Terrestrial Network (NTN). You will own 24x7 platform health across Skylo's hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). Key responsibilities include: **Cloud Infrastructure Operations & Health Ownership**: Own 24x7 cloud infrastructure health across hybrid production environments, including GKE cluster node status, on-premise Kubernetes cluster health, persistent storage operations, and infrastructure alarm triage. Execute and own Cloud Infra runbooks for P2–P4 fault categories without requiring engineering involvement for covered fault classes. **Observability Pipeline & Data Platform Operations**: Own the observability pipeline end-to-end (Prometheus, VictoriaMetrics, Grafana, OpenTelemetry), maintain database reliability (PostgreSQL streaming replication, backup/restore, failover testing), and ensure log aggregation pipeline health. **Incident Diagnosis & Escalation**: Serve as L3 escalation authority for all Cloud Infra incidents, diagnose at the Kubernetes, storage, network, and database layer, and lead troubleshooting bridges for infrastructure failures. Participate in global 24x7 on-call rotation as the Cloud Infra domain escalation tier. **SLO Engineering & Reliability**: Define and maintain SLOs for all Cloud Infra components tied to network SLA commitments to MNO partners. Own error budget tracking and drive toil reduction through automation. **Root Cause Analysis & Post-Incident Ownership**: Own Cloud Infra RCA end-to-end, deliver Initial RCA documentation within defined SLA windows, and contribute to weekly and monthly Network Performance Reports. **Runbook Authorship & Operational Standards**: Author, own, and maintain all Cloud Infra runbooks and SOPs. Define diagnostic decision trees for each known infrastructure fault class and validate operational readiness for all infrastructure changes. **Cross-Functional Collaboration**: Partner with Ops Platform Engineering, Network Implementation, Core NRE, RAN NRE, and security teams. Mentor Senior NREs in Cloud Infra domain depth. **GitOps, IaC & Platform Engineering Interface**: Own operational oversight of GitOps tooling in production (ArgoCD sync health, Helm chart version management), review Infrastructure as Code changes, and engage Platform Engineering teams with full operational context. **Requirements**: - 5+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment with direct on-call ownership for Kubernetes-at-scale environments - Deep Kubernetes expertise: multi-cluster operations (GKE or EKS), node pool management, RBAC, network policies, persistent storage (PVC, CSI drivers), CRD/operator patterns, and production cluster upgrade procedures - Hybrid cloud operations: hands-on experience operating both public cloud (GCP or AWS) and on-premise/private cloud infrastructure (bare-metal Kubernetes, KVM, or hyperconverged platforms) - Production observability stack ownership: Prometheus (federation, remote write, WAL management), Grafana, VictoriaMetrics, OpenTelemetry, and alerting pipeline design with Pub/Sub or equivalent - Database reliability: PostgreSQL streaming replication, backup/restore, failover procedures, and performance tuning; Redis cluster operations and persistence management - GitOps tooling in production: ArgoCD or Flux CD for multi-cluster operations; Helm chart authorship and version management; Terraform or Ansible for infrastructure provisioning - SRE fundamentals: SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design - Container and Linux internals: container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding - Runbook authorship: ability to write infrastructure diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure - Strong written and verbal communication: capable of delivering RCA documents, engineering escalations with structured problem statements, and MNO-facing infrastructure summaries **Preferred Qualifications**: - Experience operating cloud infrastructure for telecom or NTN workloads: 5G Core NF hosting, vRAN compute requirements, or satellite ground segment infrastructure - Software-defined storage expertise: Ceph, Rook, or equivalent distributed storage systems at production scale - Private cloud platform experience: KubeVirt, Harvester, or OpenStack for VM-container convergence on bare-metal infrastructure - Networking depth: BGP routing, VXLAN overlays, EVPN fabrics, software-defined networking, and hardware load balancer operations - Strong development background in Go or Python for building custom automation tooling, Kubernetes operators, or infrastructure lifecycle integrations - FinOps experience: cloud cost optimization, resource lifecycle automation, and capacity right-sizing across public cloud footprints - Certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect

Similar roles