SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceNow's Data Platform group is seeking a Senior Staff Engineer (IC5) to architect and deliver next-generation distributed data systems at massive scale. You will lead the design and implementation of highly scalable, high-performance platform capabilities for data-in-motion and backend storage systems, serving customers operating at the boundaries of data volume, throughput, and concurrency.
Key responsibilities include:
- Architect, design, and build high-performance distributed systems and platform components
- Build distributed systems data ingestion solutions with emphasis on scalability, quality, and operational excellence
- Deliver high-quality, modular, reusable code while enforcing engineering best practices (code reviews, unit testing, automation, design reviews)
- Build foundational libraries, frameworks, and tools focused on modularity, extensibility, configurability, and maintainability
- Collaborate across engineering teams to refine requirements and deliver end-to-end solutions
- Provide technical leadership for projects with significant complexity and risk
- Research, evaluate, and adopt new technologies that enhance platform capabilities
- Troubleshoot and diagnose complex production issues across distributed systems
This is a hands-on technical leadership role requiring deep expertise in distributed systems architecture, large-scale data platforms, and backend infrastructure. You will work with cutting-edge technologies including Kafka, Apache Iceberg, Flink, Spark, and Kubernetes to build systems that power enterprise-scale operations.
REQUIREMENTS:
- 15–20 years of hands-on engineering experience in distributed systems, data platforms, or large-scale backend infrastructure
- Strong fundamentals in distributed systems architecture, design patterns, and algorithms
- Deep programming expertise in Java, including JVM internals, memory models, and garbage collection
- Proven experience in JVM performance tuning, profiling, and diagnosing performance bottlenecks
- Strong understanding of concurrency, networking, sockets, OS internals, and performance optimization
- Hands-on experience building and operating large-scale distributed systems
- Experience with relational databases such as Oracle, MySQL, or PostgreSQL
- Deep knowledge of replication, fault-tolerant, and HA strategies
- Experience working in DevOps environments to operationalize distributed platforms
- Strong expertise in designing and architecting platforms using: Kubernetes (workload deployments, upgrades, monitoring and maintenance), Apache Iceberg (tables, catalogs, schema evolution, metadata management), Apache Kafka (high-scale clusters, topic/partition strategies, HA), Apache Flink (stateful stream processing, exactly-once semantics), Apache Spark (batch & streaming jobs, optimization, partitioning)
- Expertise in data formats such as Parquet, ORC, and Avro, along with compaction and governance strategies
- Ability to build scalable, fault-tolerant ingestion and transformation workflows
- Experience integrating Data Lakes with analytics engines, query services, or ML platforms