SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Zscaler is seeking a Senior Staff Production Engineer to join the Cloud Infrastructure & Operations team, reporting to the VP of Site Reliability Engineering. This is a hybrid role based in San Jose, CA (3 days/week) or Bellevue, WA.
You will serve as a technical leader and force multiplier for platform reliability, protecting over 15 million users globally. Your focus will be driving an "automation-first" culture across the organization while maturing observability and architectural standards to reduce Mean Time to Mitigate (MTTM) and improve scalability of Zscaler's globally distributed, multi-cloud infrastructure.
Key Responsibilities:
- Design and implement highly available, scalable infrastructure across AWS, Azure, GCP, and bare-metal environments
- Drive automation-first culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems
- Implement and maintain sophisticated observability using Prometheus, Grafana, and OpenTelemetry; define SLIs/SLOs and establish error budgets
- Act as lead Incident Commander (on-call), develop response playbooks, and conduct deep-dive post-incident analyses
- Partner with Engineering and cross-functional teams to conduct operability reviews
About Zscaler:
Zscaler is an AI-forward enterprise that accelerates digital transformation through its cloud-native Zero Trust Exchange platform. The company leverages the world's largest security data lake to protect customers from cyberattacks and data loss by securely connecting users, devices, and applications. The culture emphasizes impact over activity, constructive debate, customer obsession, collaboration, and accountability.