SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Zscaler is seeking a Director of Site Reliability Engineering to lead production engineering and SRE teams in Bangalore, reporting to the VP of Engineering. This is a hybrid role with end-to-end operational accountability for maintaining 99.99%+ availability across a cloud platform processing hundreds of billions of daily transactions.
Key Responsibilities:
- Own system architecture, automated incident mitigation, and 24/7 operational reliability for Tier-0 and Tier-1 internally facing platforms and backend infrastructure
- Enforce multi-nines SLAs, SLOs, and error budgets across services while driving engineering initiatives to reduce MTTD, MTTA, and MTTR
- Oversee capacity planning, resource forecasting, and efficiency efforts to ensure infrastructure scales reliably
- Implement high-signal alerting frameworks and actionable runbooks to minimize alert fatigue while eliminating monitoring blind spots
- Lead Production Engineering across India, driving operational rigor and coaching technical leadership on strategy, decision-making, and organizational scaling
- Partner closely with US-based leadership and globally distributed engineering teams
You will be expected to act as an owner with strong bias for action, operate with integrity, navigate seamlessly between high-level strategy and hands-on execution, and build deep empathy for internal and external customers. The role demands innovation, complex problem-solving, and the ability to lead with integrity while holding teams to high accountability standards.
Requirements:
- 12+ years of engineering experience, including 5+ years managing engineering managers and distributed teams, with proven track record as a site leader scaling and inspiring teams
- Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions within your functional domain
- Proven track record operating effectively within globally distributed engineering teams, partnering closely with US-based leadership
- Deep understanding of modern observability architectures including distributed tracing, high-cardinality metrics engines, distributed log indexing, and streaming telemetry pipelines (e.g., Kafka, OpenTelemetry, Prometheus, ClickHouse, Elasticsearch)
- Hands-on background architecting, deploying, and supporting hyper-scale cloud-native platforms running on Kubernetes and multi-cloud infrastructure
- Proven expertise in site reliability engineering principles, chaos testing, capacity planning, and automated incident triage
Preferred Qualifications:
- Proven ability to leverage AI technologies and workflows to drive measurable operational efficiency
- Bachelor's or Master's degree in Computer Science, Software Engineering, or equivalent practical experience
- Familiarity with high-volume data ingestion, backpressure management, real-time aggregation, and data lifecycle management
- Executive-level communication skills with ability to translate complex technical trade-offs for leadership, high EQ, resilience under pressure, and passion for developing people and engineering culture