SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta is seeking an experienced Senior Site Reliability Engineer to join the Emerging Products Group (EPG), focused on building highly reliable, scalable, and secure cloud services. This role combines operational excellence with software engineering, emphasizing automation-first practices and continuous improvement.
Key responsibilities include designing and operating large-scale cloud infrastructure, participating in on-call rotations for customer-facing systems, and leading incident response with post-incident reviews. You will define and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets while partnering with engineering teams to enhance availability, scalability, performance, and resilience. Observability improvements through metrics, logging, tracing, dashboards, and alerting are core to the role.
On the engineering side, you will develop software and automation using Go, Python, Terraform, and related technologies to eliminate operational toil. You'll improve deployment safety through CI/CD and GitOps practices, modernize existing workloads, and build self-service platforms and operational guardrails that enhance developer velocity while maintaining reliability and security.
Technical leadership is a key component: you will guide engineers in adopting operational best practices, mentor through technical collaboration and design reviews, support architecture decisions with data-driven recommendations, and execute projects from conception through production rollout. The role also involves exploring AI-assisted engineering techniques to improve operational efficiency and incident response.
The tech stack includes Kubernetes (EKS/GKE), Terraform, Helm, ArgoCD, Go, Python, Datadog, Splunk, PostgreSQL, Redis, and OpenSearch. This is an ideal opportunity for an experienced SRE who thrives solving complex technical challenges at scale, building automation, and improving production system reliability.