SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta is seeking a Senior Site Reliability Engineer to join its Security and Data Systems team in Bellevue, Washington. This role combines software engineering and systems administration expertise to build and maintain highly reliable, scalable, and secure infrastructure for a SaaS platform specializing in identity and security.
Key responsibilities include designing and building core infrastructure for security SaaS offerings with emphasis on high availability, performance, and scalability. You will develop robust automation using code to eliminate operational toil and ensure consistency across environments, from infrastructure provisioning to application deployment and incident response. The role requires close collaboration with security teams to embed security-first practices into all processes and infrastructure, ensuring compliance with industry standards.
You will participate in on-call rotations as a primary responder for critical incidents, leading root cause analysis and implementing preventative measures. Collaboration with development, data science, and security teams is essential to provide expert guidance on architectural decisions and best practices.
Required qualifications include U.S. Person status (U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee). You must have strong coding skills with production-level code experience, deep expertise with Terraform for infrastructure as code, and familiarity with modern CI/CD practices, particularly Spinnaker. Hands-on experience managing large-scale Kubernetes clusters, database schema management tools like Flyway, and direct experience with Snowflake data systems are essential. Experience with AI/ML applications to improve reliability and operational efficiency is a plus.
The role requires in-person onboarding and travel to the San Francisco office during the first week of employment. This is a hybrid position based in Bellevue.