SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

ServiceTitan - Remote - Remote

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ServiceTitan is seeking a Senior Site Reliability Engineer to join their Site Reliability & Infrastructure Engineering team. The role focuses on designing and maintaining the reliability and health of applications running on a cloud-native infrastructure that serves thousands of companies across the US and internationally. Key responsibilities include: - Participate in on-call rotation to diagnose and resolve production issues using runbooks and playbooks - Design, build, and maintain observability dashboards and alerting based on Service Level Indicators (SLIs) and Service Level Objectives (SLOs) - Operate and improve the Kubernetes-based compute platform that runs the majority of infrastructure - Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems - Investigate production incidents, perform root-cause analysis, and implement remediation - Partner with product engineering teams to review architecture and infrastructure decisions - Build automation to reduce manual operational work - Write and maintain runbooks and documentation to share on-call knowledge across the team - Help define non-functional requirements (scalability, availability, performance) for new systems - Collaborate across engineering teams to adopt reliability and observability best practices - Contribute to CI/CD pipelines to help teams ship changes safely and quickly Required qualifications: - Strong, hands-on understanding of Kubernetes as a system - Practical experience with SLIs, SLOs, and error budgets in real systems - Solid grounding in AWS or Azure, including networking fundamentals - Deep experience with modern observability stacks (OpenTelemetry, Prometheus, Grafana, Datadog, Elasticsearch) - Strong understanding of CI/CD systems (GitHub Actions preferred) - Strong programming skills with ability to build web applications; .NET/ASP.NET preferred, or Python/Java backgrounds acceptable - Experience with distributed systems and common failure modes - Strong production troubleshooting skills - 8-10+ years of relevant hands-on experience - Nice-to-have: database experience The ideal candidate is directly accountable for reliability of business-critical, large-scale enterprise systems, comfortable making decisions with limited information, and rewarded by developing an operability culture in a growing environment.

Similar roles