SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

ServiceNow - Santa Clara, CA, USA - Hybrid - posted 2026-09-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 149,800 - 262,200 / annual

ServiceNow is seeking a Senior Staff Software Engineer (IC5) to lead the design and delivery of shared, multi-tenant platform services on Kubernetes. You will own managed Postgres, queueing/streaming, and key-value store services that run across a large global fleet and are depended upon by product teams company-wide. In this role, you will: - Lead the design and delivery of shared platform services (managed Postgres, queueing/streaming, key-value/cache) running on Kubernetes across a global fleet, owning them from design through production operation. - Act as technical lead on major initiatives, breaking down ambiguous problems into clear, executable designs with explicit availability, durability, and cost targets. - Define HA, failover, disaster-recovery, and multi-region architecture for services you own, and build Kubernetes-native automation (operators, controllers, CRDs, self-service APIs) that provisions, scales, upgrades, and fails over systems without human intervention. - Partner with principal and distinguished engineers to align work with broader platform architecture and standards, and with product teams to set consumption contracts, tenancy models, and SLOs. - Identify technical risks early and drive resolution, with strong focus on reliability, scalability, and operability—leading failure-mode analysis, game days, and post-incident reviews for stateful systems. - Spend significant time hands-on designing, coding, and reviewing core systems such as operators, controllers, infrastructure automation, and platform services. - Mentor mid-level and junior engineers and raise the engineering bar through code reviews, design feedback, and pairing, particularly around distributed-systems and data-service design. This is a hybrid position requiring 2 days per week in a ServiceNow office (Santa Clara, San Francisco, Pleasanton, San Diego, or Kirkland). This position supports US Regulated Markets and requires passing USFedPASS screening (background check, credit check, criminal/misdemeanor check, drug test). Only US citizens, naturalized citizens, or US Permanent Residents (green card holders) will be considered due to federal requirements. REQUIREMENTS: - Experience leveraging or critically thinking about how to integrate AI into engineering and platform work—whether using AI-powered tooling, automating operational workflows, building agentic systems for fleet visibility and operations, or reasoning about AI's impact on infrastructure. - 12+ years of software development experience with a Bachelor's degree; OR 8+ years with a Master's degree; OR 5+ years with a PhD; OR equivalent work experience. - 8+ years building production software, with solid experience operating distributed systems and running stateful workloads on Kubernetes at scale. - Track record of leading the design and delivery of at least one shared, multi-tenant infrastructure service used broadly by other teams—a relational database service (Postgres or similar), a message queue or streaming platform (Kafka, NATS, RabbitMQ, or similar), or a key-value/cache service (Redis/Valkey, etcd, or similar)—including its HA, failover, and operating model. - Deep understanding of high availability and failure handling: replication topologies, leader election, quorum and consensus, split-brain avoidance, backup/restore and point-in-time recovery, and designing to explicit RPO/RTO and durability targets. - Strong system design skills—ability to reason rigorously about consistency models, partitioning and rebalancing, capacity planning, tenant isolation, and noisy-neighbor mitigation, and communicate trade-offs clearly to engineers and stakeholders. - Hands-on experience with at least one major hyperscaler (AWS, Azure, GCP), including core compute, networking, storage, and IAM primitives. - Strong working knowledge of containers and Kubernetes (including stateful primitives: StatefulSets, CSI/persistent storage, PDBs, topology spread), CI/CD and GitOps-based delivery, and infrastructure-as-code. - Strong programming skills in Go (or strong systems-language skills with willingness to work primarily in Go). NICE-TO-HAVE: - Experience building or extending Kubernetes operators that manage stateful systems (e.g., CloudNativePG, Zalando/Crunchy Postgres operators, Strimzi, Redis/Valkey operators). - Deep expertise in one of the target systems: Postgres internals (WAL, streaming/logical replication, vacuum, connection pooling, major-version upgrades); Kafka/NATS (partitioning, ISR/replication, exactly-once semantics, consumer scaling); or Redis/Valkey (cluster mode, persistence, eviction, hot-key handling). - Experience running data services across multiple regions or clusters—cross-region replication, failover orchestration, and DR testing. - Experience with zero-downtime upgrades, schema/data migrations, and fleet-wide rollouts of stateful services. - Experience with observability and SLOs for stateful systems—replication lag, saturation, tail latency, error budgets—and capacity and cost management for shared infrastructure. - Experience designing multi-tenancy: quotas, isolation, chargeback/showback, and self-service provisioning APIs or CRDs. - Experience with container networking (CNI) and/or service mesh, and with workload identity, mTLS, and secrets management as applied to data services. - Experience with managed Kubernetes (EKS/AKS/GKE), managed data services (RDS/Aurora, Cloud SQL, MSK, ElastiCache), and infrastructure-as-code tools such as Terraform or Crossplane.

Similar roles