SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 231,000 - 330,000 / annual
Zscaler is seeking an exceptional engineering executive to lead Production Engineering for its distributed Telemetry Pipelines and enterprise-wide Observability Platform. The company's cloud platform operates at immense global scale, processing hundreds of billions of transactions and petabytes of telemetry data daily. This high-visibility leadership role requires visionary systems architecture, operational rigor, and deep people leadership.
In this role, you will define the long-term roadmap and architectural strategy for the global Observability Platform, driving strategic consolidation and modernization of monitoring stacks on open standards like OpenTelemetry to optimize total cost of ownership. You will architect, scale, and optimize high-throughput, low-latency telemetry ingestion and processing pipelines handling petabyte-scale streaming data to ensure fault tolerance, durability, and cost efficiency. You will optimize platform performance and resource utilization to maintain industry-leading service margins while handling aggressive year-over-year telemetry volume growth.
You will partner with Site Reliability Engineering, Core Product, and incident management teams to establish standardized SLOs, enforce error budgets, and ensure high-fidelity alerting with streamlined escalation workflows. You will lead, mentor, and inspire a world-class engineering organization across multiple global sites while fostering an inclusive, accountable culture focused on technical curiosity and zero-toil automation. You will act as a trusted partner to Product Management, Information Security, and Customer Support leadership to align operational telemetry with customer business outcomes.
The role is based in San Jose, CA and requires hybrid work (three days per week on-site), reporting to executive engineering leadership.
REQUIREMENTS:
- Experience leveraging generative AI tools, AI/ML models, or intelligent agents to optimize production engineering workflows, automate incident response toil, or enhance predictive observability analytics
- Progressive engineering leadership experience totaling 7+ years managing managers, technical leads, and distributed engineering organizations of 30+ engineers within hyper-scale SaaS, cloud provider, or high-transaction distributed systems environments
- Deep subject matter expertise in modern observability and telemetry systems, including distributed tracing, high-cardinality metrics engines, distributed log analytics, and telemetry routing using technologies such as OpenTelemetry, Kafka, Flink, Vector, VictoriaMetrics, Grafana, Prometheus, ClickHouse, or OpenSearch
- Strong architectural foundation in cloud-native paradigms, container orchestration with Kubernetes, and microservices, combined with practical grasp of Site Reliability Engineering methodologies and incident command structures
- Proven track record of influencing executive stakeholders and driving cross-functional alignment, with ability to work on-site at headquarters at least 3 days per week
PREFERRED QUALIFICATIONS:
- Advanced experience with compliance, data retention, PII masking, and data governance requirements within large-scale logging and telemetry environments
- Active contribution to open-source communities in the observability or telemetry space, such as OpenTelemetry or Prometheus
- Experience designing and implementing chaos engineering practices or automated disaster recovery testing at scale