SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceTitan is seeking a Principal Site Reliability Engineer to join their Site Reliability & Infrastructure Engineering team. This is a technical leadership role where you'll define how the organization thinks about reliability across all engineering teams, not just operate the platform.
You'll set the technical direction and establish standards for infrastructure and reliability practices. Key responsibilities include designing and delivering major platform initiatives around Kubernetes infrastructure, observability frameworks, SLO programs, and incident management systems. You'll partner with Engineering Managers, Staff Engineers, and product leadership to review architecture decisions and maintain non-functional requirements across the organization.
You'll identify systemic reliability risks before they become incidents and drive remediation across teams. A major focus is defining and operationalizing SLIs, SLOs, and error budgets as organization-wide standards, not just within the SRE team. You'll drive adoption of reliability and observability best practices through documentation, design reviews, and direct partnerships with product teams.
The role includes leveraging AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to automate investigations and accelerate incident resolution. You'll mentor Staff and Senior SREs, raising the technical ceiling of the team through code reviews, architecture feedback, and complex problem-solving partnerships. You'll also contribute to technical hiring and help define Principal-level expectations at ServiceTitan.
Required expertise includes deep, expert-level understanding of Kubernetes internals, failure modes, and large-scale cluster management. You need proven experience defining and operationalizing SLIs, SLOs, and error budgets across engineering organizations. Deep expertise with modern observability stacks (OpenTelemetry, Prometheus, Grafana, Datadog, Elasticsearch) and ability to design org-wide observability frameworks is essential.
You should have expert grounding in AWS, GCP, or Azure including networking, security, and cost optimization at scale. Distributed systems depth is required—ability to reason through complex failure modes and architect graceful degradation. Track record of leading major incident response, driving blameless postmortems, and implementing systemic fixes is expected. Deep CI/CD experience and hands-on experience with AI-assisted SRE capabilities in production systems are required.
The ideal candidate has 12+ years of relevant hands-on experience with a clear track record of technical leadership on complex, cross-team infrastructure initiatives at product companies operating at scale. You've operated at Staff or Principal level before and know what it means to own a technical domain. You're the person Engineering Managers and VPs call when a system needs to be rethought. You can set direction with incomplete information, make architectural trade-offs explicit, and build consensus across teams without organizational authority.