SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer - Splunk

Okta - San Francisco, CA, United States - Hybrid - posted 2026-02-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 194,000 - 267,000 / annual

Okta is seeking a Staff Site Reliability Engineer specializing in Splunk observability to own and evolve the company's observability ecosystem. This is a technical individual contributor role focused on building a world-class, scalable observability platform that serves SRE teams and business partners across the organization. Key responsibilities include designing and maintaining scalable observability infrastructure using Terraform and infrastructure-as-code principles; optimizing Splunk's collection, processing, and storage of log data for high reliability and low latency; participating in on-call rotations and leading post-incident reviews to drive systemic improvements; and automating deployment and scaling of observability agents and collectors to eliminate operational toil. Required qualifications include 5+ years of experience scaling and managing Splunk Cloud at enterprise scale (1000+ services), with expertise in Workload Management (WLM) and HEC optimization. You must have 5+ years in an SRE, DevOps, or Systems Engineering role focused on high-availability systems. Strong programming proficiency in SPL and Go is essential for building internal tools and automating workflows. Deep understanding of Linux internals, networking (TCP/IP, DNS, load balancing), and container orchestration (Kubernetes/EKS) is required. You should demonstrate a data-driven approach to debugging complex, cross-service performance bottlenecks and expertise in creating intuitive, actionable Splunk dashboards that correlate data across multiple sources. Bonus skills include hands-on experience with OpenTelemetry (OTel), Vector, or similar telemetry frameworks; experience implementing Splunk charge-back applications for usage reporting; and experience managing observability tools within AWS or GCP. Note: This role requires U.S. Person status (U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee) to access federal environments and protected federal data. In-person onboarding in the San Francisco office is required during the first week of employment.

Similar roles