SlipstreamJobsFresh Startup & VC-Backed Jobs

Manager, Site Reliability Engineering (Auth0)

Okta - New York, NY, United States - Hybrid - posted 2026-08-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Okta is seeking a Manager of Site Reliability Engineering to lead the SRE team at Auth0, a critical authentication platform serving millions of users globally. This is a hands-on leadership role combining strategic direction with deep technical involvement. You will lead the SRE team's technical strategy, translating organizational vision into actionable roadmaps and driving complex cross-functional initiatives across product and platform teams. The role requires active participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems. Key responsibilities include building infrastructure resilience through monitoring, alerting, and automation improvements; championing reliability best practices and embedding observability into all engineering efforts; mentoring and developing SRE talent through pair programming and design discussions; and representing reliability as a senior technical leader in architectural reviews and strategic planning. You bring 3+ years of hands-on team leadership in SRE or software engineering within cloud-native environments, plus 8+ years total industry experience. Deep expertise in AWS/Azure, infrastructure-as-code (Terraform), and cloud-native architectures (containers, Kubernetes, microservices, databases) is essential. Strong programming skills in Go or Python, with a track record of building production-grade tools and automation, are required. You operate with a data-driven mindset grounded in SRE principles: blameless culture, systematic problem-solving, and applying software engineering to operational challenges. Exceptional communication skills—verbal and written—enable you to drive clarity during high-pressure incidents and articulate complex concepts to diverse stakeholders. You have proven ability to build and lead high-performing teams in globally distributed, remote-first environments. Extra credit for experience leading reliability initiatives that improved system uptime and reduced incident response times at scale, contributions to open-source infrastructure or observability tooling, and experience designing comprehensive incident response programs and runbook automation. Note: This position requires ability to access federal environments and establish U.S. Person status upon hire.

Similar roles