SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: EUR 64,000 - 88,000 / annual
Auth0, part of Okta, is seeking a Senior Site Reliability Engineer to join its Europe-based SRE team. You'll be responsible for designing and building custom software in Go to enhance platform reliability, resiliency, and redundancy for a service that handles authentication for hundreds of millions of users worldwide.
Key responsibilities include partnering with engineering teams to embed reliability principles across services, improving availability, performance, and observability. You'll leverage deep infrastructure and observability expertise to identify and implement improvements, contribute to a follow-the-sun on-call rotation (with shifts only during your local working hours), and develop SRE tooling and processes focused on automation and operational efficiency. You'll also define and champion reliability best practices across the organization.
Successful candidates will have proven experience supporting large-scale, mission-critical applications in production environments with high autonomy. You should be proficient in at least one programming language (Go preferred), with hands-on experience in infrastructure as code (Terraform), container orchestration (Kubernetes, Docker), and GitOps (ArgoCD). Expertise with major cloud providers (Azure, AWS, or GCP) is required, along with strong understanding of microservices architecture, databases (SQL and NoSQL), and networking fundamentals.
You'll need demonstrated knowledge of core SRE principles including SLIs, SLOs, and error budgets, plus experience in 24/7 on-call rotations for cloud-based environments. A proactive, systematic approach to problem-solving with high ownership mentality is essential. Exceptional communication and collaboration skills are critical for working effectively in a remote, distributed team where tasks are often self-driven. This is a hands-on builder role focused on directly contributing to platform resiliency at scale.