SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

Replit - Remote - Remote - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Replit is an agentic software creation platform that democratizes application development by enabling anyone to build using natural language. The company serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will be a senior technical leader responsible for ensuring the reliability, scalability, and performance of Replit's infrastructure. You will bridge development and operations, designing and implementing systems that enable the platform to scale efficiently while maintaining high availability. Key responsibilities include: - Architect and implement comprehensive observability solutions (monitoring, logging, tracing) with real-time dashboards and metrics for proactive issue detection. - Define and drive Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across product and engineering teams, building systems to track and report on these metrics. - Lead high-impact incident response as a senior leader, conduct blameless post-mortems, and drive preventative measures. Develop runbooks and automation to reduce Mean Time To Recovery (MTTR). - Design and implement infrastructure automation using tools like Terraform or Pulumi to eliminate operational toil. Architect self-healing systems that respond automatically to common failure scenarios. - Optimize performance across large-scale cloud deployments with deep focus on Kubernetes, Docker, and GCP. Identify bottlenecks, implement capacity planning, and reduce latency across global regions. - Debug complex distributed systems issues and design long-term fixes that improve robustness, operability, and diagnosability. - Review feature and system designs company-wide as a key owner for reliability, scalability, security, and operational integrity. - Mentor and educate the broader engineering team to embed reliability as a core cultural value. - Write high-quality, well-tested code in Python or Go for internal tools and third-party integrations. REQUIREMENTS: - 8-10 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Infrastructure Engineering. - Strong programming skills in Python or Go with ability to write high-quality, well-tested code. - Deep understanding of distributed systems design, production service scaling, and service-oriented architecture. - Deep experience with Kubernetes and cloud-native container orchestration platforms. - Proven track record designing, implementing, and maintaining sophisticated monitoring and observability solutions (metrics, logging, tracing). - Strong incident management skills with extensive experience leading incident response for complex systems and critical thinking under pressure. - Experience with infrastructure as code (Terraform, Pulumi) and configuration management tools. - Excellent written and verbal communication skills; ability to explain complex technical concepts clearly. - Strong interpersonal skills with experience mentoring engineers from junior to principal levels. - Willingness to dive into and improve any layer of the technical stack. - Passion for making software creation accessible and empowering developers. BONUS: - Deep experience with Google Cloud Platform (GCP) services and tools. - Expert-level knowledge of modern observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry). - Experience designing reliable systems for high throughput and low latency. - Significant Go and Terraform experience. - Familiarity with rapid-growth startup environments. - Experience writing company-facing blog posts and training materials.

Similar roles