SlipstreamJobsFresh Startup & VC-Backed Jobs

Sr. Staff Site Reliability Engineer

Zscaler - Bangalore, KA, India - Hybrid - posted 2026-09-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Zscaler is seeking a Senior Staff Site Reliability Engineer to join the Production Engineering and Reliability team in Bangalore. You will own the reliability and operational integrity of a global infrastructure fleet spanning GCP, AWS, and bare-metal systems, building self-healing automation, hardening provisioning workflows, and embedding observability across the platform. Key responsibilities include: - Develop production-grade automation in Ansible, Python, and/or Go for bare-metal provisioning, IPMI/Redfish-based hardware lifecycle operations, and autonomous remediation - Design and implement fault-tolerant provisioning systems with self-healing and auto-remediation capabilities - Develop telemetry pipelines for hardware and system monitoring, integrating metrics, logs, and distributed tracing (Prometheus, Grafana, Victoria Metrics, Telegraf) - Collaborate with cross-functional teams to deliver integrated solutions and contribute to Agile/Scrum processes - Partner directly with software engineering teams to define and enforce production-readiness criteria, including SLOs, health-check standards, and rollback requirements You will report to the Director of Engineering in the Production Engineering and Reliability team. This role requires someone who thrives in ambiguity, acts with ownership, solves hard problems, collaborates at high trust, and maintains a growth mindset. You are passionate about the mission and operate with integrity, seeing challenges as opportunities to deliver meaningful impact. Zscaler is a NASDAQ-listed company pioneering the Zero Trust Exchange platform, protecting thousands of customers from cyberattacks and data loss. The company is building an AI-native enterprise where human potential is amplified by machine intelligence. Requirements: - 7+ years of Production Engineering, Platform Engineering, or Infrastructure Engineering experience in cloud or hybrid environments - Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions within your functional domain - Deep Linux/Unix systems mastery including OS internals, kernel networking, boot pipelines, and low-level system debugging - Hands-on operational experience across GCP and/or AWS alongside on-premises datacenter operations - Working command of infrastructure-as-code tooling (Terraform, Ansible) for both cloud and on-prem environments - Expertise building durable, idempotent, and repeatable automation using Python and/or Go Preferred qualifications: - Experience implementing AIOps, AI-driven predictive scaling, or automated anomaly detection to optimize cloud infrastructure performance and reliability - Experience building or operating stateful orchestration systems using Temporal.io, Cadence, or equivalent - Built and scaled a full observability stack from scratch

Similar roles