SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Okta's TDI Network Engineering team is seeking a Staff Reliability Engineer to design, own, and ensure the resilience, health, and availability of the company's global corporate network. This operations-focused role reports to the Network Engineering Manager and centers on operational execution—responding to alerts, maintaining service availability, and ensuring system health across the enterprise network.
Key responsibilities include:
- Designing and owning the resilience and availability of the entire global corporate network domain, managing operational responsibilities such as alert response, health monitoring, and executing reliability projects to maintain an "Always Secure. Always On." environment.
- Driving strategic reduction of systemic toil and technical debt across multiple teams by introducing process efficiencies, automating network operations, and building scalable self-service operational tooling.
- Collaborating with cross-functional stakeholders including Business Technology, Workplace, Security, and executive leaders, challenging assumptions and cascading relevant information to project teams.
- Making critical decisions and leading resolution of complex network operations issues from a systems perspective, anticipating business challenges and preventing future outages.
- Fostering learning and talent development by defining success for the team, cultivating transparency, and actively mentoring team members through the P4 level.
You will own multi-quarter objectives and establish long-term strategies for network reliability, applying deep expertise in Distributed Systems, Networking fundamentals, Infrastructure as Code, and observability to architect scalable platforms.
Requirements:
- Typically 8+ years of related professional experience with a Bachelor's degree; or 6+ years with a Master's degree; or 3+ years with a PhD; or equivalent experience.
- Deep expertise in AWS Networking and Palo Alto Networks solutions as core required technical competencies.
- Comprehensive operational experience in Distributed Systems and Networking fundamentals, including quick incident response, system health monitoring, and management of WiFi, DNS, DHCP, VLANs, VPN, ACLs, Routing, and Firewall Policies.
- Strong proficiency in Cloud Platforms, Infrastructure as Code (Terraform/Ansible), Observability tools (Prometheus/Grafana), Programming (Python/Go), and Service Reliability Management (SLOs/SLIs) for large-scale enterprise environments.
- Proven track record of managing operational availability, delivering multi-quarter objectives, and executing technical projects within defined budgets and strategic targets.
- Demonstrated ability to navigate high levels of ambiguity, establish credibility with executive stakeholders, and model resilience during major system transitions or production incidents.
- Desirable: Experience with Juniper/JUNOS switching/routing, Palo Alto Networks NGFWs, and enterprise office build and construction processes.