SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Replit is an agentic software creation platform that enables users to build applications using natural language. This Engineering Manager role leads the Site Reliability Engineering function across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure.
You will lead and grow an existing SRE team that builds and operates production platforms, working hands-on across application and infrastructure boundaries. This is a software-building leadership role focused on helping teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements.
Key responsibilities include:
- Building and operating metrics, logs, traces, and alerting capabilities; helping teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.
- Owning incident tooling and practices, coordinating cross-team response, and turning incident reviews into engineering improvements that reduce recovery time and repeat failures.
- Building and maintaining load/failure testing capabilities to validate critical paths under expected demand and test recovery and production readiness.
- Leading deep engagements with internal teams on SLOs and end-to-end performance, using profiling, telemetry, and load tests to identify bottlenecks and deliver improvements.
- Staying technically engaged by reviewing designs and production changes, debugging failure modes, and using AI coding tools to prototype and automate.
- Building and growing a high-ownership engineering team through coaching, developing technical leaders, managing performance, and hiring.
- Measuring outcomes by tracking rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements from cost/capacity analysis.
Requirements:
- Demonstrated engineering management experience: you have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team.
- Software-oriented production systems depth: you have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.
- Safe-change and performance judgment: you have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems.
- Platform-product and cross-team judgment: you can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.
Nice to have:
- Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.
- Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.
- Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.
- Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.