SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Replit - Remote - Remote - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Replit is an agentic software creation platform that democratizes application development by enabling anyone to build applications using natural language. The company serves millions of developers worldwide. As a Senior Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and performance of Replit's infrastructure. You will bridge development and operations, implementing automation and establishing best practices that enable the platform to scale efficiently while maintaining high availability. Key responsibilities include: • Design and implement comprehensive observability solutions using modern monitoring and alerting tools. Create dashboards and metrics providing real-time visibility into system health and performance. Implement logging strategies enabling quick problem identification and resolution. • Architect and implement infrastructure automation using tools like Terraform, Ansible, or Pulumi. Design and maintain CI/CD pipelines enabling reliable and consistent deployments. Create self-healing systems that automatically respond to common failure scenarios. • Work with product and engineering teams to define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to track and report on these metrics, ensuring high reliability standards while balancing innovation speed. • Lead incident response efforts, conduct thorough post-mortems, and implement improvements to prevent future occurrences. Develop and maintain runbooks for critical services. Build tools and processes that reduce Mean Time To Recovery (MTTR). • Identify and resolve performance bottlenecks across infrastructure. Implement capacity planning strategies and optimize resource utilization. Work on reducing latency and improving system efficiency across global regions. The company values problem-solving mindset, self-directed autonomous work, strong communication skills, continuous learning, and a focus on automation. You will work in an autonomous environment with flexible time off, competitive benefits including equity, health insurance, paid leave, and quarterly team gatherings. REQUIREMENTS: • 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering) • Strong programming skills in languages commonly used for automation (Python, Go, or similar) • Deep understanding of distributed systems • Experience with container orchestration platforms (Kubernetes) and cloud-native technologies • Proven track record of implementing and maintaining monitoring/observability solutions • Strong incident management skills with experience leading incident response • Experience with infrastructure as code and configuration management tools BONUS: • Experience with Google Cloud Platform (GCP) services and tools • Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.)

Similar roles