SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta is seeking a Staff Site Reliability Engineer to design, build, and manage Kubernetes platforms that power the Workforce Identity Cloud. This is a hands-on technical role focused on architecting highly available, scalable, and secure cloud-native infrastructure on AWS.
You will own the full lifecycle of Kubernetes platform creation and optimization, including cluster design, multi-region deployments, and production hardening. Key responsibilities include managing AWS infrastructure (EKS, ECS, S3, RDS, VPCs, IAM), automating deployments via Helm charts, implementing dynamic scaling with Karpenter, and configuring Istio service mesh for secure service-to-service communication. You'll design and implement CI/CD pipelines, automate infrastructure provisioning with Terraform, and ensure observability through monitoring and logging tools like Prometheus, Grafana, and CloudWatch.
Incident response and troubleshooting are core to the role—you'll diagnose and resolve production issues related to performance, availability, and security. You'll also drive platform automation, cost optimization, and security best practices across multi-region cloud environments. Documentation and knowledge sharing are expected to promote operational excellence across teams.
This role requires deep technical expertise in Kubernetes, AWS, and infrastructure-as-code, with a proven track record of building and scaling production systems. You should be comfortable with scripting (Python, Bash, Go) and have hands-on experience with modern DevOps tooling. The ideal candidate combines strong technical depth with the ability to mentor others and drive architectural decisions that impact the entire platform.