SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

Okta - Dublin, Ireland - Hybrid - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: EUR 92,000 - 126,500 / annual

Okta is seeking a Staff Site Reliability Engineer to join its Dublin/EMEA infrastructure team. You will be the senior-most technical voice on the ground, setting architectural direction for the Kubernetes platform, driving reliability and scaling challenges to resolution, and partnering with Staff/Principal engineers across US and India teams to maintain global platform strategy coherence. You will report to the Senior Manager, Site Reliability Engineering. Key responsibilities include: - Owning the architecture and evolution of the Kubernetes platform for Dublin/EMEA, including cluster strategy, multi-tenancy, networking, and security posture in coordination with global counterparts - Leading technical design and rollout of self-service platform capabilities, golden paths, and self-healing patterns to enable internal teams to ship and operate workflows without deep platform expertise - Driving platform readiness for AI-driven internal workflows at scale, including model serving, GPU scheduling, and orchestration infrastructure - Taking ownership of the hardest, highest-ambiguity reliability problems: complex production incidents, capacity and scaling challenges, and cross-system failure modes - Participating in follow-the-sun on-call rotation with SRE peers across regions, driving root-cause resolution and postmortems - Setting technical standards for SLIs/SLOs, error budgets, and on-call practice, holding the team accountable - Improving SDLC processes for infrastructure-as-code, including CI/CD pipeline maturity and change/release management - Mentoring engineers across the team and raising the technical bar through design reviews, architecture discussions, and hands-on pairing - Partnering with architects, security, and product engineering teams to align infrastructure decisions with business reliability, security, and delivery needs - Representing the Dublin/EMEA team in global platform architecture discussions as a peer with Staff/Principal-level engineers Requirements: - 8+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering with a track record of staff-level technical leadership - Deep, hands-on expertise in Kubernetes at production scale (architecture, multi-tenancy, networking, security, day-2 operations) - Experience running large-scale internal infrastructure platforms in public cloud, preferably AWS - Strong expertise in cloud-native architectures, infrastructure-as-code (Terraform), and CI/CD pipelines - Experience building or operating infrastructure that hosts AI/ML or agentic workflows (model serving, GPU scheduling, orchestration frameworks) - Deep experience with observability platforms and monitoring tools (Grafana, Splunk, APM, or equivalent) in large-scale environments - Demonstrated ability to influence technical direction across teams and geographies without formal authority - Effective verbal and written communication skills with experience mentoring engineers and driving technical alignment across distributed teams - Computer Science degree or related field, or equivalent experience Extra credit for: internal developer platform (IDP) experience with platform-as-product mindset, AI agent orchestration/vector databases/model routing/inference optimization, multi-cloud experience (AWS plus Azure/GCP), service mesh technologies (Istio, Linkerd) or GitOps tooling (ArgoCD, Flux) in production Kubernetes, Kubernetes certifications (CKA, CKS, CKAD) or equivalent cloud certifications, enterprise-scale SCM platform management experience (GitHub, GitLab).

Similar roles