SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta is seeking a Staff Site Reliability Engineer to architect and evolve the company's global core network infrastructure and multi-cloud platform layers. This role is foundational to building highly resilient, secure systems that support Okta's Workforce Identity Cloud platform.
Key responsibilities include designing next-generation global ingress and edge infrastructure using Cloudflare for SaaS, eliminating legacy hardware routing bottlenecks, and shifting compute to edge workers to reduce latency for millions of global requests. You will build and scale Okta's multi-cloud backbone across AWS and GCP, architecting hub-and-spoke models using AWS Transit Gateway, AWS Global Accelerator, and GCP Network Connectivity Center.
The role emphasizes zero-trust security, requiring implementation and management of strict cryptographic controls at scale, including TLS 1.3 externally and deep mTLS mesh architectures internally across microservice proxies (Envoy/Nginx). You will innovate on platform operations by designing deterministic playbooks and automated telemetry ingestion pipelines using AI models and LLMs, moving the team from reactive debugging to autonomous anomaly detection and self-healing systems.
Additional responsibilities include providing first-line technical triage during high-volume DDoS attacks during India Data Center hours, analyzing real-time proxy and VPC flow logs, and implementing rapid WAF mitigations. You will champion Infrastructure-as-Code principles using Terraform, continuously identifying and eliminating network configuration drifts across cloud landing zones. The role includes mentoring a lean 3-member local team while collaborating with US-based engineering leaders, and creating detailed global architecture blueprints, technical documentation, and disaster recovery runbooks.
Required qualifications include 8+ years of production cloud infrastructure engineering experience, 5+ years with enterprise cloud networking constructs (AWS Transit Gateway, VPC peering, AWS Global Accelerator, Direct Connect, GCP equivalents), strong knowledge of edge delivery networks and edge compute stacks (Cloudflare for SaaS, Cloudflare Tunnels, Edge Workers), and deep understanding of network layers and security protocols including HTTP/S routing, TCP/IP stack debugging, PKI, certificate rotation, and proxy layers. Advanced proficiency with Terraform and Infrastructure-as-Code tools is essential.