SlipstreamJobsFresh Startup & VC-Backed Jobs

Director, Data & Storage Reliability Engineering

ServiceNow - Santa Clara, CA, United States - In-office - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ServiceNow is seeking a Director of Data & Storage Reliability Engineering to lead a large, multi-disciplinary engineering organization focused on improving reliability, resilience, performance, scalability, and observability across the company's database, storage, and platform infrastructure. In this role, you will build and scale high-performing engineering teams spanning reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering. You will own talent acquisition, performance management, career development, succession planning, and strategic objective setting while establishing a strong engineering-first culture centered on data-driven decision making and operational excellence. Key responsibilities include identifying recurring failure patterns, reliability risks, performance bottlenecks, and scalability constraints across database services, storage platforms, and cloud infrastructure. You will drive engineering improvements that eliminate entire classes of issues before they impact customers. You will partner closely with Product Engineering, Database Engineering, Cloud Infrastructure, Architecture, Storage Engineering, Support, and Operations teams to embed reliability considerations throughout the software development lifecycle. You will serve as the senior technical leader for complex reliability investigations, customer-critical escalation reviews, and migration readiness assessments. You will influence architectural decisions and technology investments by providing reliability expertise, observability insights, and production-based evidence that improve platform resilience and customer outcomes. The role requires a strong product mindset—you will treat reliability, observability, resilience, and automation capabilities as products with roadmaps, priorities, adoption goals, and measurable outcomes. You will establish scalable reliability engineering practices, standards, governance processes, and operating models across the organization. You will drive adoption of observability standards, reliability frameworks, resiliency assessments, and automation strategies. You will establish meaningful KPIs and engineering metrics providing visibility into platform reliability, operational efficiency, customer experience, and risk reduction. You will leverage AI-powered tools, analytics, and production intelligence to identify emerging risks, improve detection coverage, and accelerate engineering insights. You will champion a proactive reliability engineering model that shifts the organization from reactive issue response toward predictive analysis, prevention, and continuous optimization.

Similar roles