SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta is seeking an experienced Staff Site Reliability Engineer to join the Emerging Products Group (EPG), where you'll serve as a technical leader driving reliability, scalability, and operational excellence across cloud services. This role combines hands-on engineering with technical leadership, focusing on building highly reliable, secure cloud infrastructure that customers can trust.
You will design, build, and operate large-scale cloud infrastructure and production services, participating in on-call rotations for highly available customer-facing systems. You'll lead incident response efforts and drive post-incident reviews focused on systemic improvements, while defining and improving Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partnering with engineering teams, you'll enhance service availability, scalability, performance, and resilience, and continuously improve observability through metrics, logging, tracing, dashboards, and alerting.
On the engineering side, you'll develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies. You'll eliminate operational toil through automation, tooling, and platform engineering, improve deployment safety through CI/CD and GitOps practices, and build self-service platforms and operational guardrails that improve developer velocity while maintaining reliability and security.
As a technical leader, you'll guide complex reliability initiatives spanning multiple engineering teams, mentor engineers in operational best practices and reliability engineering principles, influence architecture and operational decisions through data-driven recommendations, and drive projects from conception through production rollout. You'll also explore AI-assisted engineering techniques to improve operational efficiency, incident response, and troubleshooting.
The tech stack includes Kubernetes (EKS/GKE), Terraform, Helm, Git, ArgoCD, GitOps, Golang, Python, Datadog, Splunk, PostgreSQL, Redis, and OpenSearch. You'll need strong experience operating large-scale production services in AWS and/or GCP, with a passion for automation, continuous improvement, and operational excellence.