SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 221,200 - 387,100 / annual
ServiceNow is seeking a Director of Site Reliability Engineering & Service Enablement to lead the next phase of the company's reliability transformation as it modernizes toward a cloud-agnostic, cloud-ready production platform.
You will own key elements of the SRE operating model across Reliability Engineering, Service Enablement, Service Registry, SLI/SLO standards, reliability governance, automation, AI-enabled operations, and production readiness. You will lead a global engineering organization and partner across Product Engineering, Infrastructure, Architecture, Security, Release Engineering, and Customer Support to establish consistent reliability practices across ServiceNow products and services.
Key responsibilities include:
- Define and execute the SRE strategy and operating model across reliability engineering, service enablement, observability, automation, incident learning, and production readiness
- Lead and develop a global organization of engineering managers, technical leaders, and SREs
- Establish enterprise reliability standards for service ownership, tiering, golden signals, SLIs/SLOs, error budgets, alerting, on-call practices, and service health reviews
- Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services
- Own the Service Registry strategy, improving service ownership, dependency visibility, maturity tracking, and impact-aware operational decision-making
- Drive adoption of SLIs, SLOs, error budgets, and burn-rate alerting across critical services
- Build a culture of engineering away toil by converting recurring operational work and incident patterns into automation, self-service, and systemic fixes
- Establish the AI-enabled SRE roadmap, including change-risk assessment, operational insights, remediation recommendations, and policy-driven automation
- Drive reliability and production-readiness strategy across AWS, Azure, and GCP by establishing cloud-agnostic patterns
- Partner with product and platform engineers to design, launch, and operate reliable services throughout the production lifecycle
- Establish launch and production-readiness practices that validate availability, latency, performance, capacity, dependencies, rollback, and recovery
- Drive sustainable operations by scaling self-service capabilities, automation platforms, and systemic reliability improvements
- Lead incident response, blameless postmortems, and corrective actions that convert production failures into lasting reliability improvements
- Measure reliability through SLIs, SLOs, error budgets, golden signals, change failure rate, MTTR, capacity health, and toil reduction
- Influence architecture and platform direction to simplify operating models and improve reliability
- Partner with executive and engineering leaders to prioritize reliability investments and drive adoption beyond the direct SRE organization
You will play a critical role in evolving the organization from reactive operations toward an engineering-led SRE model focused on prevention, automation, resilience, and continuous improvement.
REQUIREMENTS:
- 12 years of significant leadership experience in Site Reliability Engineering, Production Engineering, Platform Engineering, Cloud Infrastructure, or large-scale distributed systems with a Bachelor's degree; OR 8 years and a Master's degree; OR a PhD with 5 years experience; OR equivalent experience
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving
- Proven success leading managers and senior technical leaders across geographically distributed engineering organizations
- Demonstrated success leading SRE, infrastructure, or reliability transformation at scale
- Strong understanding of SLIs/SLOs, error budgets, observability, incident management, reliability governance, and on-call practices
- Experience with service catalogs, service registries, service ownership models, Backstage, CMDB, dependency mapping, or service topology
- Strong background in cloud infrastructure and modernization across AWS, Azure, and/or GCP
- Understanding of Kubernetes, distributed systems, networking, databases, infrastructure automation, and cloud-native architecture
- Experience driving automation through orchestration, Infrastructure as Code, self-service platforms, and auto-remediation
- Familiarity with AI-assisted operations, autonomous remediation, or agentic technologies is highly desirable
- Experience establishing production-readiness practices for releases, resilience, disaster recovery, infrastructure changes, and cloud migrations
- Ability to use incident, reliability, and operational data to prioritize engineering work and drive systemic improvements
- Strong cross-functional influence and executive communication skills
- Ability to operate effectively through ambiguity, organizational transformation, and large-scale technical change