SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceNow is seeking a Senior Staff Software Engineer to lead the design and operation of cloud-native reliability, release, and test platforms that enable engineering excellence and high-confidence releases through automation, observability, and AI-driven operations.
In this role, you will:
• Build and operate cloud-native engineering platforms for software validation, release qualification, and operational readiness
• Design production-like release and test environments that improve release confidence and deployment readiness
• Develop automated quality gates to assess release health, operational risk, and production readiness
• Integrate automated testing, observability, reliability signals, and deployment intelligence into CI/CD pipelines
• Build reusable test frameworks, self-service environments, test data, mock services, and developer productivity tooling
• Advance shift-left engineering through automated validation, continuous verification, and quality gates
• Automate failure detection, policy validation, deployment verification, security checks, and reliability assessments
• Lead Kubernetes-based platform evolution for scalable test infrastructure, release automation, and developer self-service
• Resolve recurring infrastructure issues through sustainable software, systems, and networking solutions
• Partner with engineering teams on design reviews, architecture standards, and automation-first reliability practices
This is a platform engineering and SRE role focused on building systems that improve developer productivity and release confidence at enterprise scale. You will work on mission-critical infrastructure supporting ServiceNow's AI control tower platform used by 85% of the Fortune 500.
REQUIREMENTS:
• 12+ years of experience in software, systems, platform, or reliability engineering
• Deep Kubernetes expertise across architecture, operations, networking, storage, security, autoscaling, and multi-cluster environments
• Experience building and operating large-scale Kubernetes platforms for cloud-native, mission-critical services
• Experience integrating Kubernetes with CI/CD, GitOps, automated testing, and deployment validation
• Experience designing cloud-native platforms for ephemeral environments, release qualifications, and automated validation
• Proven ability to lead engineering excellence across developer productivity, platform engineering, release confidence, and modernization
• Experience with progressive delivery, including canary releases, feature flags, automated rollback, and deployment verification
• Experience with chaos engineering, resilience validation, disaster recovery, and reliability assessments
• Expertise designing, authoring, testing, and debugging code in a team setting using languages such as Python, Go, Java, or Ruby
• Experience using AI-assisted engineering for intelligent testing, release risk analysis, incident diagnostics, and operational automation
• Strong coding, observability, SLO, and cross-team collaboration skills to improve reliability, performance, and engineering standards
GOOD TO HAVE:
• Expertise in observability and monitoring applications, services, and networks at scale
• Experience with DevOps automation, CI/CD pipelines, and agile methodologies, including GitLab CI/CD or similar tools
• Experience building enterprise-scale test automation frameworks such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG
• Experience with test orchestration, test impact analysis, flaky test detection, parallel execution, and intelligent regression testing
• Experience with service virtualization, contract testing, synthetic testing, and test data management
• Experience building engineering platforms that support developer self-service and release engineering
• Experience with infrastructure configuration management tools such as Ansible
• Expertise with Kubernetes ecosystem technologies such as Helm, Argo CD, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtimes
• Experience implementing GitOps using Argo CD, Flux, or similar technologies
• Experience operating Kubernetes across AWS (EKS), Azure (AKS), and Google Cloud (GKE)