SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineering Manager

Shippo - Remote - Remote - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 175,000 - 238,000 / annual

Shippo is building the shipping layer of the internet, providing logistics technology and infrastructure that connects merchants to carriers worldwide through APIs and dashboards. As a remote-first, globally distributed company, Shippo serves the e-commerce backbone. You will lead a team of platform-focused SRE engineers at Shippo, responsible for building platforms, tooling, and infrastructure that enable product teams to operate reliable, performant, and scalable services. Your team will establish frameworks for observability, deployment automation, and infrastructure management, allowing product teams to own their service reliability while you maintain a support-oriented culture focused on automation and operational excellence. Key responsibilities include: - Leading and developing a team of SRE engineers with technical mentorship, career development, and performance management - Building and maintaining internal platforms and tooling for service deployment, monitoring, and operations - Managing observability platforms (metrics, logs, traces, dashboards) providing visibility into services - Owning the infrastructure and Kubernetes platform that all Shippo services run on, with capacity planning and performance optimization - Establishing SLO/SLI frameworks and error budget tracking for product teams - Designing and maintaining CI/CD pipelines, deployment automation, and release tooling - Building infrastructure-as-code foundations and self-service capabilities - Creating automation to eliminate toil and prevent infrastructure problems - Driving infrastructure cost optimization initiatives - Participating in Sev1 incident leadership rotation - Managing the SRE team's on-call rotation - Designing, implementing, and testing disaster recovery capabilities - Ensuring infrastructure security and compliance - Partnering with Engineering Managers and TPMs on platform roadmap and capabilities - Establishing platform SLOs for infrastructure reliability, deployment success, and developer experience metrics Requirements: - 3+ years of hands-on engineering management experience - 9+ years as a software or systems engineer with deep experience building platforms, tooling, or infrastructure - BS or MS degree in Computer Science or equivalent experience - Expert-level experience designing and operating platforms that enable other engineering teams (internal platform-as-a-product) - Strong operational experience with Kubernetes in production environments, including building Kubernetes platforms for application teams - Deep expertise with at least one public cloud provider (AWS, GCP) including networking, compute, storage, and managed services - Experience building or maintaining CI/CD systems and deployment automation (GitHub Actions, GitLab CI, ArgoCD, Flux, etc.) - Strong background in infrastructure-as-code tools and patterns (Terraform, Pulumi, CloudFormation, etc.) - Experience designing and implementing observability platforms (Prometheus, Grafana, ELK stack, Datadog, New Relic, etc.) - Proficiency in at least one programming language for tooling and automation (Python, Go, or similar) - Experience establishing reliability frameworks (SLO/SLI/error budgets) that other teams can adopt - Understanding of developer experience and ability to build self-service tooling that reduces friction - Track record of designing disaster recovery capabilities

Similar roles