SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 221,200 - 387,100 / annual
ServiceNow is seeking a Director of Data & Storage Reliability Engineering to lead a strategic organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across the company's database, storage, and platform infrastructure.
In this role, you will build, develop, and scale high-performing engineering teams focused on reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering. You will be accountable for talent acquisition, performance management, career development, succession planning, objective setting, coaching, and prioritization of strategic initiatives.
Key responsibilities include:
- Establishing a strong engineering-first culture centered on data-driven decision making, continuous improvement, operational excellence, and systemic risk reduction
- Identifying recurring failure patterns, reliability risks, performance bottlenecks, and scalability constraints across database services, storage platforms, and cloud infrastructure
- Driving engineering improvements that eliminate entire classes of issues before they impact customers
- Partnering with Product Engineering, Database Engineering, Cloud Infrastructure, Architecture, Storage Engineering, Support, and Operations teams to ensure reliability and performance considerations are embedded throughout the software development lifecycle
- Establishing a continuous feedback loop between production operations and platform improvement initiatives
- Serving as the senior technical leader for complex reliability investigations, customer-critical escalation reviews, and migration readiness assessments
- Influencing architectural decisions and technology investments through reliability expertise and production-based evidence
- Treating reliability, observability, resilience, performance, and automation capabilities as products with roadmaps, priorities, adoption goals, and measurable outcomes
- Establishing scalable reliability engineering practices, standards, governance processes, and operating models
- Driving adoption of observability standards, reliability engineering frameworks, resiliency assessments, and automation strategies
- Establishing meaningful KPIs and engineering metrics that provide visibility into platform reliability, operational efficiency, and customer experience
- Leveraging AI-powered tools, analytics, and automation frameworks to identify emerging risks and improve engineering productivity
- Maintaining a portfolio of reliability investments balancing immediate customer needs with long-term platform strategy
- Championing a proactive reliability engineering model that shifts from reactive issue response toward predictive analysis and prevention
Requirements:
- 15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environments
- 8+ years of engineering leadership experience, including leading managers and globally distributed teams
- Extensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizations
- Deep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architectures
- Strong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineering
- Experience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scale
- Proven experience identifying systemic issues and converting operational insights into strategic engineering improvements
- Strong product mindset with demonstrated experience treating technical capabilities as products with roadmaps, priorities, customers, adoption goals, and measurable business outcomes
- Experience translating production insights, customer pain points, and operational challenges into prioritized engineering investments and long-term roadmaps
- Experience partnering closely with production operations, customer escalation teams, and software engineering teams to drive systemic improvements
- Experience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations without direct authority
- Experience operating a portfolio of engineering investments, balancing short-term customer needs with long-term reliability and resilience objectives
- Experience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomes
- Experience driving engineering initiatives through data, metrics, customer impact analysis, and measurable business outcomes
- Experience leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving
- Experience leveraging AI technologies to improve decision-making, analytics, engineering workflows, operational efficiency, reliability insights, automation, or customer outcomes
- Exceptional communication, stakeholder management, and leadership skills
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
Desired skills include previous Product Management experience in platform, infrastructure, cloud, database, or SaaS environments; experience operating large-scale enterprise database and storage platforms; experience building and scaling Reliability Engineering or SRE organizations; experience with observability platforms and telemetry systems; experience with migration readiness and large-scale cloud transformations; experience with Linux-based production environments; experience supporting enterprise database technologies such as MySQL, PostgreSQL, Oracle, or SQL Server; and familiarity with ServiceNow platform architecture.