SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Okta - Dublin, Ireland - Hybrid - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: EUR 76,000 - 104,500 / annual

Okta is seeking a Senior Site Reliability Engineer to join its Dublin/EMEA team within TDI Infrastructure Engineering. You will build and operate the Kubernetes-based platform that hosts internal workflows across the business, with a growing focus on hosting AI-driven workflows reliably in production. You will report to the Senior Manager, Site Reliability Engineering, and work closely with a Staff SRE and the Dublin team, as well as SRE peers in the US and India. Key responsibilities include: - Building, operating, and improving the Kubernetes platform — cluster management, networking, security posture, and day-2 operations - Developing self-service capabilities, golden paths, and automation that let internal teams onboard and operate workflows with less platform expertise required - Supporting the platform's readiness to host AI-driven internal workflows reliably at scale - Troubleshooting and resolving production issues, including participating in a follow-the-sun on-call rotation shared with SRE peers globally, and driving postmortems for incidents you own - Contributing to and upholding reliability practices — SLIs/SLOs, error budgets, and operational runbooks - Improving CI/CD pipelines and infrastructure-as-code practices, including change and release management - Building and maintaining observability and monitoring coverage for the systems you own - Mentoring junior engineers and contributing to design and code reviews - Collaborating with the Staff SRE and Dublin team to align local implementation with the broader global platform strategy Okta authenticates, authorizes, and provisions millions of users a day. The company is focused on securing AI by building trusted, neutral infrastructure that enables organizations to safely embrace this new era. REQUIREMENTS: - 5+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering - Strong hands-on experience with Kubernetes in production — deployment, networking, security, and troubleshooting - Experience operating large-scale internal infrastructure platforms in a public cloud, preferably AWS - Solid experience with infrastructure-as-code (Terraform) and CI/CD pipelines - Exposure to or interest in building infrastructure that hosts AI/ML or agentic workflows - Experience with observability platforms and monitoring tools (Grafana, Splunk, APM, or equivalent) - Comfortable working in a fast-moving, evolving environment with ambiguity around process and tooling - Effective verbal and written communication skills, with the ability to collaborate across time zones - Computer Science degree or related field, or equivalent experience Extra credit for: internal developer platform (IDP) concepts or "platform as a product" thinking; hands-on experience with AI agent orchestration, vector databases, model routing, or inference infrastructure; multi-cloud experience (AWS plus Azure or GCP); service mesh technologies (Istio, Linkerd) or GitOps tooling (ArgoCD, Flux); Kubernetes certifications (CKA, CKS, CKAD) or equivalent cloud certifications; experience administering an enterprise-scale SCM platform (GitHub, GitLab, or equivalent).

Similar roles