SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 125,800 - 211,600 / annual
Zillow Group is seeking a Senior Manager (M4 level) to lead the incident and problem management function across the organization. This role owns the end-to-end strategy, processes, and people behind incident response and problem management, driving operational excellence and reliability across production systems.
Key responsibilities include:
Incident Management Leadership: Own the incident management program including process design, tooling, and governance. Serve as executive escalation point for critical incidents. Set standards for severity classification, escalation paths, and communication protocols. Partner with Engineering, Product, and business leadership to align incident response with business priorities.
Problem Management: Build and own a formal problem management practice connecting incident trends to systemic root causes. Ensure significant incidents produce rigorous, blameless root cause analyses with clearly owned corrective actions. Identify recurring issues and chronic risks. Hold cross-functional partners accountable for remediation timelines. Report on outcomes and reliability trends to leadership.
AI-Enabled Operations: Champion adoption of AI-powered tooling across incident detection, triage, summarization, and RCA drafting. Design workflows that turn raw incident data into actionable insights. Ensure AI-generated outputs meet high accuracy standards. Identify opportunities to automate repetitive tasks. Partner with Engineering and Data teams to pilot and scale new AI capabilities.
People & Team Leadership: Hire, coach, and develop a team of incident managers. Set clear performance expectations and career growth paths. Establish on-call structures and rotations that sustain team health. Foster a culture of ownership, continuous improvement, and blameless learning.
Operational Excellence: Define and track key metrics (MTTR, MTTD, recurrence rate, action-item closure rate). Continuously refine runbooks and workflows based on retrospectives. Facilitate post-incident reviews for major incidents. Benchmark practices against industry standards.
The role directly manages a team of incident managers and indirectly influences engineering and support teams during incident response. Decisions and process changes shape reliability posture and stakeholder trust across the organization.