SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Staff Reliability Engineer

ServiceNow - Santa Clara, CA, United States - Hybrid - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 190,900 - 334,100 / annual

ServiceNow is seeking a Senior Staff Reliability Engineer to design, build, and operate cloud-native engineering platforms that enable software validation, release automation, and production readiness through advanced automation, observability, and AI-driven operations. In this role, you will: - Design and build cloud-native engineering platforms for software validation, release validation, and production readiness - Design and maintain production-like release and test environments that improve release confidence and deployment readiness - Build and integrate automated test pipelines, observability, reliability signals, deployment intelligence, and quality gates into CI/CD workflows - Develop automation solutions that improve engineering productivity and reduce manual toil through shift-left engineering practices - Build reusable frameworks, self-service engineering environments, test data management, mock services, and developer productivity tooling - Design and enhance Kubernetes-based platforms supporting scalable test infrastructure, release automation, cloud-native workloads, and developer self-service - Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, resilience testing, and operational health assessments - Resolve complex platforms, infrastructure, and networking challenges through software engineering, systems design, and automation - Partner closely with engineering teams to improve platform reliability, release quality, cloud-native adoption, and engineering best practices - Participate in architecture reviews and technical design discussions for scalable, automation-first engineering solutions - Influence technical decisions through strong engineering execution, collaboration, and delivery of high-quality platform capabilities - Mentor engineers through technical guidance, code reviews, knowledge sharing, and engineering best practices - Foster a culture of reliability, automation, operational excellence, continuous improvement, and customer-focused engineering Requirements: - 12+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree; OR 8+ years with a Master's degree; OR a PhD with 5+ years experience; OR equivalent experience - Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments - Experience building and operating cloud-native platforms supporting scalable, highly available services - Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud-native deployment workflows - Experience designing and implementing automation to improve developer productivity, release quality, and operational efficiency - Experience with progressive delivery practices, including canary deployments, feature flags, automated rollback, and deployment verification - Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation - Strong software engineering skills with hands-on experience designing, developing, testing, and debugging applications using Python, Go, Java, or Ruby - Strong understanding of observability, monitoring, SLI/SLOs, incident management, and production operations for distributed systems - Demonstrated ability to solve complex technical problems, drive projects independently, and collaborate effectively across engineering teams - Experience leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving - Thrives in fast-paced, ambiguous environments with strong ownership mindset, bias for action, and passion for continuous learning and automation - Low ego, intellectually curious, and effective collaborator who enjoys partnering with globally distributed teams Plus (nice to have): - Experience with observability and monitoring platforms for applications, services, and distributed systems at scale - Experience with DevOps automation, CI/CD pipelines, GitOps, and Agile development practices using tools such as GitLab CI/CD, Argo CD, or Flux - Experience building and maintaining enterprise-scale test automation frameworks using technologies such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG - Experience with test orchestration, intelligent regression testing, test impact analysis, flaky test detection, parallel execution, and test data management - Experience with service virtualization, contract testing, synthetic testing, and building developer self-service engineering platforms - Experience with Infrastructure as Code and configuration management tools such as Ansible, Terraform - Experience with Kubernetes ecosystem tools including Helm, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry - Experience operating Kubernetes platforms across public cloud providers (AWS EKS, Azure AKS, Google Cloud GKE) - Familiarity with AI-assisted engineering, intelligent testing, operational automation, or cloud-native engineering platforms

Similar roles