SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta is seeking a Staff Site Reliability Engineer to architect and manage Kubernetes platforms supporting cloud-native applications at scale. This role is central to Okta's Workforce Identity Cloud, which enables secure access for enterprise workforces.
You will design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms on AWS. Key responsibilities include:
• Kubernetes Platform Creation: Design and operate production-grade Kubernetes clusters optimized for resilience and operational efficiency.
• AWS Infrastructure Management: Build and optimize AWS infrastructure (EKS, ECS, S3, VPCs, RDS, IAM) with focus on cost management, scaling, and security.
• Helm & Karpenter: Automate application deployments using Helm charts and implement dynamic cluster scaling with Karpenter.
• Istio Service Mesh: Configure and manage Istio for service-to-service communication, security, and observability.
• Platform Automation: Develop CI/CD pipelines and infrastructure-as-code (Terraform, Ansible) to automate deployment and scaling.
• Incident Management: Respond to and resolve system issues related to performance, availability, and security.
• Security & Compliance: Design secure cloud infrastructure with appropriate access controls, network security, and compliance frameworks.
• Documentation & Knowledge Sharing: Create operational procedures and promote best practices across teams.
Required: 4+ years Kubernetes/Helm, 4+ years Terraform, 5+ years AWS, multi-region cloud experience, strong Kubernetes platform expertise, Helm deployment skills, Karpenter optimization, Istio service mesh management, CI/CD pipeline proficiency (Jenkins, GitLab, CircleCI, Spinnaker), scripting in Python/Bash/Go, and monitoring tools (Prometheus, Grafana, CloudWatch, ELK Stack).
Preferred: Cloud security best practices, Docker/containerization knowledge, Bachelor's in Computer Science or equivalent.