SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer, Federal (TS/SCI)

Okta - Washington, DC, United States - In-office - posted 2026-07-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Okta is seeking an experienced Staff Site Reliability Engineer to join the Federal SRE team within the Emerging Products Group (EPG). This role is ideal for an engineer who thrives solving complex technical challenges at scale, building automation, and improving production system reliability. You will serve as a technical leader within the EPG SRE organization, partnering with software engineers, architects, and product teams to design, build, and operate world-class cloud services. Key responsibilities include designing and operating large-scale cloud infrastructure and production services; participating in on-call rotations for highly available customer-facing systems; leading incident response efforts and driving post-incident reviews focused on systemic improvements; and defining, measuring, and improving Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. You will partner with engineering teams to improve service availability, scalability, performance, and resilience, and continuously improve observability through metrics, logging, tracing, dashboards, and alerting. On the engineering and automation side, you will develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies; eliminate operational toil through automation, tooling, and platform engineering; improve deployment safety through CI/CD and GitOps practices; and build self-service platforms and operational guardrails that improve developer velocity while maintaining reliability and security. As a technical leader, you will lead complex reliability initiatives spanning multiple engineering teams, guide engineers in adopting operational best practices and reliability engineering principles, mentor engineers through technical collaboration and design reviews, influence architecture decisions through data-driven recommendations, and drive projects from conception through production rollout. Required qualifications include an active U.S. TS/SCI security clearance with Full Scope Poly and proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6). You should have strong experience operating large-scale production services in AWS and/or GCP, deep expertise with Kubernetes in production environments, extensive experience with Infrastructure as Code technologies such as Terraform and Helm, strong software engineering skills in Golang and/or Python, and experience building automation and internal engineering platforms. Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra is essential, along with strong understanding of cloud networking fundamentals.

Similar roles