SlipstreamJobsFresh Startup & VC-Backed Jobs

Director, Site Reliability Engineering & Service Enablement

ServiceNow - Santa Clara, CA, United States - Hybrid - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 221,200 - 387,100 / annual

ServiceNow is seeking a Director of Site Reliability Engineering & Service Enablement to lead the next phase of the company's reliability transformation as it modernizes toward a cloud-agnostic, cloud-ready production platform. You will own key elements of the SRE operating model across Reliability Engineering, Service Enablement, Service Registry, SLI/SLO standards, reliability governance, automation, AI-enabled operations, and production readiness. You will lead a global engineering organization and partner across Product Engineering, Infrastructure, Architecture, Security, Release Engineering, and Customer Support to establish consistent reliability practices across ServiceNow products and services. Key responsibilities include: - Define and execute the SRE strategy and operating model across reliability engineering, service enablement, observability, automation, incident learning, and production readiness - Lead and develop a global organization of engineering managers, technical leaders, and SREs - Establish enterprise reliability standards for service ownership, tiering, golden signals, SLIs/SLOs, error budgets, alerting, on-call practices, and service health reviews - Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services - Own the Service Registry strategy, improving service ownership, dependency visibility, maturity tracking, and impact-aware operational decision-making - Drive adoption of SLIs, SLOs, error budgets, and burn-rate alerting across critical services - Build a culture of engineering away toil by converting recurring operational work and incident patterns into automation, self-service, and systemic fixes - Establish the AI-enabled SRE roadmap, including change-risk assessment, operational insights, remediation recommendations, and policy-driven automation - Drive reliability and production-readiness strategy across AWS, Azure, and GCP by establishing cloud-agnostic patterns - Partner with product and platform engineers to design, launch, and operate reliable services throughout the production lifecycle - Establish launch and production-readiness practices that validate availability, latency, performance, capacity, dependencies, rollback, and recovery - Drive sustainable operations by scaling self-service capabilities, automation platforms, and systemic reliability improvements - Lead incident response, blameless postmortems, and corrective actions that convert production failures into lasting reliability improvements - Measure reliability through SLIs, SLOs, error budgets, golden signals, change failure rate, MTTR, capacity health, and toil reduction - Influence architecture and platform direction to simplify operating models and improve reliability - Partner with executive and engineering leaders to prioritize reliability investments and drive adoption beyond the direct SRE organization You will play a critical role in evolving the organization from reactive operations toward an engineering-led SRE model focused on prevention, automation, resilience, and continuous improvement. REQUIREMENTS: - 12 years of significant leadership experience in Site Reliability Engineering, Production Engineering, Platform Engineering, Cloud Infrastructure, or large-scale distributed systems with a Bachelor's degree; OR 8 years and a Master's degree; OR a PhD with 5 years experience; OR equivalent experience - Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving - Proven success leading managers and senior technical leaders across geographically distributed engineering organizations - Demonstrated success leading SRE, infrastructure, or reliability transformation at scale - Strong understanding of SLIs/SLOs, error budgets, observability, incident management, reliability governance, and on-call practices - Experience with service catalogs, service registries, service ownership models, Backstage, CMDB, dependency mapping, or service topology - Strong background in cloud infrastructure and modernization across AWS, Azure, and/or GCP - Understanding of Kubernetes, distributed systems, networking, databases, infrastructure automation, and cloud-native architecture - Experience driving automation through orchestration, Infrastructure as Code, self-service platforms, and auto-remediation - Familiarity with AI-assisted operations, autonomous remediation, or agentic technologies is highly desirable - Experience establishing production-readiness practices for releases, resilience, disaster recovery, infrastructure changes, and cloud migrations - Ability to use incident, reliability, and operational data to prioritize engineering work and drive systemic improvements - Strong cross-functional influence and executive communication skills - Ability to operate effectively through ambiguity, organizational transformation, and large-scale technical change

Similar roles