SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta is hiring a Staff Site Reliability Engineer to architect and manage Kubernetes platforms supporting cloud-native applications on AWS. This is a high-impact role within the Workforce Identity Cloud team, focused on building reliable, scalable, and secure infrastructure for production workloads.
Key responsibilities include designing and maintaining highly available Kubernetes clusters optimized for production; managing AWS infrastructure (EKS, ECS, RDS, S3, IAM, VPCs); creating and maintaining Helm charts for automated deployments; implementing Karpenter for dynamic cluster scaling; configuring Istio service mesh for service-to-service communication and observability; automating deployment and scaling workflows with CI/CD pipelines; responding to and resolving incidents; implementing security and compliance controls; and documenting platform procedures and best practices.
Required qualifications: 5+ years with Kubernetes, Helm, Karpenter, and Istio; 8+ years with infrastructure-as-code tools (Terraform, Chef, Ansible); 8+ years with serverless computing (AWS Lambda, API Gateway) and microservices architecture; proven AWS expertise (EKS, ECS, RDS, S3, CloudFormation, IAM); strong Kubernetes platform creation and optimization skills; hands-on Helm deployment experience; practical Karpenter implementation; Istio service mesh expertise; CI/CD pipeline proficiency (Jenkins, GitLab, CircleCI, Spinnaker); and strong scripting/automation skills in Python or Go.
The ideal candidate exemplifies the principle "if you have to do something more than once, automate it" and can rapidly self-educate on new concepts and tools. This role offers the opportunity to solve large-scale automation, testing, and tuning problems while contributing to Okta's mission of securing AI and identity infrastructure.