SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceTitan is seeking a Senior Site Reliability Engineer to join their Site Reliability & Infrastructure Engineering team. The role focuses on designing and maintaining the reliability and health of applications running on a cloud-native infrastructure that serves thousands of companies across the US and internationally.
Key responsibilities include:
- Participate in on-call rotation to diagnose and resolve production issues using runbooks and playbooks
- Design, build, and maintain observability dashboards and alerting based on Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- Operate and improve the Kubernetes-based compute platform that runs the majority of infrastructure
- Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems
- Investigate production incidents, perform root-cause analysis, and implement remediation
- Partner with product engineering teams to review architecture and infrastructure decisions
- Build automation to reduce manual operational work
- Write and maintain runbooks and documentation to share on-call knowledge across the team
- Help define non-functional requirements (scalability, availability, performance) for new systems
- Collaborate across engineering teams to adopt reliability and observability best practices
- Contribute to CI/CD pipelines to help teams ship changes safely and quickly
Required qualifications:
- Strong, hands-on understanding of Kubernetes as a system
- Practical experience with SLIs, SLOs, and error budgets in real systems
- Solid grounding in AWS or Azure, including networking fundamentals
- Deep experience with modern observability stacks (OpenTelemetry, Prometheus, Grafana, Datadog, Elasticsearch)
- Strong understanding of CI/CD systems (GitHub Actions preferred)
- Strong programming skills with ability to build web applications; .NET/ASP.NET preferred, or Python/Java backgrounds acceptable
- Experience with distributed systems and common failure modes
- Strong production troubleshooting skills
- 8-10+ years of relevant hands-on experience
- Nice-to-have: database experience
The ideal candidate is directly accountable for reliability of business-critical, large-scale enterprise systems, comfortable making decisions with limited information, and rewarded by developing an operability culture in a growing environment.