SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

ServiceNow - Dublin, Ireland - Hybrid - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ServiceNow is seeking a Staff Site Reliability Engineer to design, build, and operate cloud-native engineering platforms that support software validation, release validation, and production readiness. This is a high-impact technical role where you'll architect and maintain production-like environments, build automated test pipelines with observability and quality gates, and develop automation solutions that reduce manual toil through shift-left engineering practices. Key responsibilities include: - Design and build Kubernetes-based platforms supporting scalable test infrastructure, release automation, and developer self-service - Integrate automated test pipelines, observability, reliability signals, and deployment intelligence into CI/CD workflows - Develop reusable frameworks, self-service engineering environments, and developer productivity tooling - Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, and resilience testing - Resolve complex platform, infrastructure, and networking challenges through software engineering and systems design - Partner with engineering teams to improve platform reliability, release quality, and cloud-native adoption - Participate in architecture reviews and technical design discussions - Mentor engineers through technical guidance, code reviews, and knowledge sharing - Foster a culture of reliability, automation, and continuous improvement This role emphasizes technical leadership and influence through strong engineering execution, collaboration, and delivery of high-quality platform capabilities. You'll work on mission-critical infrastructure that impacts the entire engineering organization. Requirements: - 8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering (with Bachelor's degree); or 6 years with Master's degree; or PhD with 3+ years experience; or equivalent - Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments - Experience building and operating cloud-native platforms supporting scalable, highly available services - Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, and cloud-native deployment workflows - Experience designing and implementing automation to improve developer productivity, release quality, and operational efficiency - Experience with progressive delivery practices (canary deployments, feature flags, automated rollback, deployment verification) - Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation - Strong software engineering skills with hands-on experience in Python, Go, Java, or Ruby - Strong understanding of observability, monitoring, SLI/SLOs, incident management, and production operations for distributed systems - Demonstrated ability to solve complex technical problems, drive projects independently, and collaborate effectively - Ownership mindset, bias for action, passion for continuous learning and automation - Experience leveraging or critically thinking about how to integrate AI into work processes and problem-solving Preferred qualifications: - Experience with observability and monitoring platforms at scale - DevOps automation, CI/CD pipelines, GitOps using GitLab CI/CD, Argo CD, or Flux - Enterprise-scale test automation frameworks (Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG) - Test orchestration, intelligent regression testing, test impact analysis, flaky test detection - Service virtualization, contract testing, synthetic testing - Infrastructure as Code tools (Ansible, Terraform) - Kubernetes ecosystem tools (Helm, Argo Workflows, Kustomize, Istio/Linkerd, Prometheus, OpenTelemetry) - Operating Kubernetes on AWS (EKS), Azure (AKS), Google Cloud (GKE) - AI-assisted engineering and intelligent testing

Similar roles