SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

Okta - Bengaluru, India - In-office - posted 2026-07-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Okta is seeking an experienced Staff Site Reliability Engineer to join the Emerging Products Group (EPG), where you'll serve as a technical leader driving reliability, scalability, and operational excellence across cloud services. This role combines hands-on engineering with technical leadership—you'll design and operate large-scale cloud infrastructure, lead incident response efforts, and mentor engineers across the organization. Key responsibilities include: **Reliability & Operations**: Design and operate large-scale cloud infrastructure supporting highly available customer-facing systems. Participate in on-call rotations, lead incident response efforts, and drive post-incident reviews focused on systemic improvements. Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously enhance observability through metrics, logging, tracing, dashboards, and alerting. **Engineering & Automation**: Develop software and infrastructure using Go, Python, Terraform, and related technologies. Eliminate operational toil through automation, tooling, and platform engineering. Improve deployment safety through CI/CD and GitOps practices. Build self-service platforms and operational guardrails that improve developer velocity while maintaining reliability and security. **Technical Leadership**: Lead complex reliability initiatives spanning multiple engineering teams. Guide engineers in adopting operational best practices and reliability engineering principles. Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing. Influence architecture and operational decisions through data-driven recommendations. Drive projects from conception through production rollout. **Innovation**: Explore AI-assisted engineering techniques to improve operational efficiency, incident response, and troubleshooting. Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity. Tech stack includes Kubernetes (EKS/GKE), Terraform, Helm, ArgoCD, Golang, Python, Datadog, Splunk, PostgreSQL, Redis, and OpenSearch.

Similar roles