SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Zscaler is seeking a Senior Production Engineer to join the Cloud Infrastructure & Operations team, reporting to the VP of Site Reliability Engineering. This is a hybrid role (3 days/week in San Jose, CA) with flexibility for Bellevue, WA or remote US options.
You will be a force multiplier for the reliability of Zscaler's globally distributed, multi-cloud platform that protects over 15 million users. Your focus will be on providing technical vision and hands-on execution to drive an "automation-first" culture across the organization.
Key Responsibilities:
- Design and implement highly available, scalable infrastructure across AWS, Azure, GCP, and bare-metal environments
- Drive automation-first culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems
- Implement and maintain sophisticated observability using Prometheus, Grafana, and OpenTelemetry; define SLIs/SLOs and establish error budgets
- Act as a lead Incident Commander (on-call), develop response playbooks, and conduct deep-dive post-incident analyses
- Partner with Engineering and cross-functional teams to conduct operability reviews
- Directly reduce Mean Time to Mitigate (MTTM) and shape the scalability of globally distributed infrastructure
You thrive in ambiguity and act like an owner with a bias for action. You're a problem-solver who loves running toward challenges, a high-trust collaborator ambitious for the team, and a continuous learner with a growth mindset. You operate with integrity and navigate seamlessly between high-level strategy and hands-on execution.