SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: EUR 76,000 - 104,500 / annual
Okta is seeking a Senior Site Reliability Engineer to join its Dublin/EMEA team within TDI Infrastructure Engineering. You will build and operate the Kubernetes-based platform that hosts internal workflows across the business, with a growing focus on hosting AI-driven workflows reliably in production. You will report to the Senior Manager, Site Reliability Engineering, and work closely with a Staff SRE and the Dublin team, as well as SRE peers in the US and India.
Key responsibilities include:
- Building, operating, and improving the Kubernetes platform — cluster management, networking, security posture, and day-2 operations
- Developing self-service capabilities, golden paths, and automation that let internal teams onboard and operate workflows with less platform expertise required
- Supporting the platform's readiness to host AI-driven internal workflows reliably at scale
- Troubleshooting and resolving production issues, including participating in a follow-the-sun on-call rotation shared with SRE peers globally, and driving postmortems for incidents you own
- Contributing to and upholding reliability practices — SLIs/SLOs, error budgets, and operational runbooks
- Improving CI/CD pipelines and infrastructure-as-code practices, including change and release management
- Building and maintaining observability and monitoring coverage for the systems you own
- Mentoring junior engineers and contributing to design and code reviews
- Collaborating with the Staff SRE and Dublin team to align local implementation with the broader global platform strategy
Okta authenticates, authorizes, and provisions millions of users a day. The company is focused on securing AI by building trusted, neutral infrastructure that enables organizations to safely embrace this new era.
REQUIREMENTS:
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering
- Strong hands-on experience with Kubernetes in production — deployment, networking, security, and troubleshooting
- Experience operating large-scale internal infrastructure platforms in a public cloud, preferably AWS
- Solid experience with infrastructure-as-code (Terraform) and CI/CD pipelines
- Exposure to or interest in building infrastructure that hosts AI/ML or agentic workflows
- Experience with observability platforms and monitoring tools (Grafana, Splunk, APM, or equivalent)
- Comfortable working in a fast-moving, evolving environment with ambiguity around process and tooling
- Effective verbal and written communication skills, with the ability to collaborate across time zones
- Computer Science degree or related field, or equivalent experience
Extra credit for: internal developer platform (IDP) concepts or "platform as a product" thinking; hands-on experience with AI agent orchestration, vector databases, model routing, or inference infrastructure; multi-cloud experience (AWS plus Azure or GCP); service mesh technologies (Istio, Linkerd) or GitOps tooling (ArgoCD, Flux); Kubernetes certifications (CKA, CKS, CKAD) or equivalent cloud certifications; experience administering an enterprise-scale SCM platform (GitHub, GitLab, or equivalent).