SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

ServiceNow - San Diego, CA, United States - Hybrid - posted 2026-09-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 149,800 - 262,200 / annual

ServiceNow is seeking a Senior Staff Software Engineer (IC5 level) to join the Data Platform Engineering organization. This is a hands-on technical leadership role focused on designing and delivering shared, multi-tenant platform services—managed Postgres, queueing/streaming, and key-value stores—running on Kubernetes across a large global fleet. In this role, you will: - Lead the design and delivery of shared platform services that product teams across the company depend on, owning them from design through production operation - Act as the technical lead on major initiatives, breaking down ambiguous problems into clear, executable designs with explicit availability, durability, and cost targets - Define HA, failover, disaster-recovery, and multi-region architecture for owned services, and build Kubernetes-native automation (operators, controllers, CRDs, self-service APIs) that provisions, scales, upgrades, and fails over systems without human intervention - Partner with principal and distinguished engineers to align work with broader platform architecture and standards, and with product teams to set consumption contracts, tenancy models, and SLOs - Identify technical risks early and drive resolution, with strong focus on reliability, scalability, and operability—leading failure-mode analysis, game days, and post-incident reviews for stateful systems - Spend significant time hands-on designing, coding, and reviewing core systems such as operators, controllers, infrastructure automation, and platform services - Mentor mid-level and junior engineers and raise the engineering bar through code reviews, design feedback, and pairing, particularly around distributed-systems and data-service design This is a hybrid position requiring 2 days per week in a ServiceNow office. Offices are located in San Francisco, Pleasanton, Santa Clara, San Diego, and Kirkland. This position supports US Regulated Markets and requires passing USFedPASS (US Federal Personnel Authorization Screening Standards), including background check, credit check, criminal/misdemeanor check, and drug test. Due to federal requirements, only US citizens, US naturalized citizens, or US Permanent Residents holding a green card will be considered. REQUIREMENTS: Must-have experience: - 12+ years of software development with a Bachelor's degree; OR 8+ years with a Master's degree; OR 5+ years with a PhD; OR equivalent work experience - 8+ years building production software with solid experience operating distributed systems and running stateful workloads on Kubernetes at scale - Track record of leading the design and delivery of at least one shared, multi-tenant infrastructure service used broadly by other teams—such as a relational database service (Postgres or similar), message queue or streaming platform (Kafka, NATS, RabbitMQ, or similar), or key-value/cache service (Redis/Valkey, etcd, or similar)—including its HA, failover, and operating model - Deep understanding of high availability and failure handling: replication topologies, leader election, quorum and consensus, split-brain avoidance, backup/restore and point-in-time recovery, and designing to explicit RPO/RTO and durability targets - Strong system design skills—ability to reason rigorously about consistency models, partitioning and rebalancing, capacity planning, tenant isolation, and noisy-neighbor mitigation, and communicate trade-offs clearly to engineers and stakeholders - Hands-on experience with at least one major hyperscaler (AWS, Azure, GCP), including core compute, networking, storage, and IAM primitives - Strong working knowledge of containers and Kubernetes (including stateful primitives: StatefulSets, CSI/persistent storage, PDBs, topology spread), CI/CD and GitOps-based delivery, and infrastructure-as-code - Strong programming skills in Go (or strong systems-language skills with willingness to work primarily in Go) - Experience leveraging or critically thinking about how to integrate AI into engineering and platform work—whether using AI-powered tooling, automating operational workflows, building agentic systems for fleet visibility and operations, or reasoning about AI's impact on how infrastructure is built and run Nice-to-have experience: - Building or extending Kubernetes operators that manage stateful systems (e.g., CloudNativePG, Zalando/Crunchy Postgres operators, Strimzi, Redis/Valkey operators) - Deep expertise in Postgres internals (WAL, streaming/logical replication, vacuum, connection pooling with PgBouncer/PgCat, major-version upgrades); Kafka/NATS (partitioning, ISR/replication, exactly-once semantics, consumer scaling); or Redis/Valkey (cluster mode, persistence, eviction, hot-key handling) - Running data services across multiple regions or clusters—cross-region replication, failover orchestration, and DR testing - Zero-downtime upgrades, schema/data migrations, and fleet-wide rollouts of stateful services - Observability and SLOs for stateful systems—replication lag, saturation, tail latency, error budgets—and capacity and cost management for shared infrastructure - Designing multi-tenancy: quotas, isolation, chargeback/showback, and self-service provisioning APIs or CRDs - Container networking (CNI) and/or service mesh, and workload identity, mTLS, and secrets management as applied to data services - Managed Kubernetes (EKS/AKS/GKE), managed data services (RDS/Aurora, Cloud SQL, MSK, ElastiCache), and infrastructure-as-code tools such as Terraform or Crossplane

Similar roles