SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 232,000 - 319,000 / annual
Okta is seeking a Senior Manager of Site Reliability Engineering to lead the Infrastructure Platform and Shared Services organization. You will oversee multiple technical teams focused on edge networking, Kubernetes platforms, observability, automation, and tooling that support Okta's identity and access management service, which authenticates millions of users daily across AWS infrastructure with 99.999% availability.
In this role, you will lead and mentor a high-performing team of engineers and managers, driving initiatives across the SRE and infrastructure organization. Key responsibilities include building a world-class observability platform with self-service capabilities, accelerating product engineering velocity through robust platforms and intuitive tooling, and owning the design and operation of scalable cloud infrastructure platforms. You will perform engineering design evaluations, improve SDLC processes for infrastructure-as-code, manage service expectations, and prioritize resource allocation across multiple domains.
You bring 6+ years of technical leadership and people management experience, with 3+ years running large-scale infrastructure platforms supporting SaaS/cloud services on AWS (multi-cloud experience is a plus). You have strong expertise in cloud-native architectures, edge infrastructure (WAF, ALB, NLB, Apache, Nginx), infrastructure-as-code (Terraform), observability tools (Splunk, Grafana, APM), and CI/CD pipelines. You have demonstrated hands-on experience with SRE automation and tooling, a track record leading cross-functional teams on large-scale programs, and strong communication skills. A computer science degree or equivalent experience is required.
Note: This position requires U.S. Person status and the ability to access federal environments as a condition of employment.