SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer - Infrastructure,Kubernetes

ServiceNow - Hyderabad, India - Hybrid - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ServiceNow is hiring a Staff Software Engineer (IC4) for the Telemetry Data Platform team, which handles product usage telemetry across the entire ServiceNow platform. The team operates a distributed infrastructure stack built on Kubernetes, Kafka, and ClickHouse, with team members across Israel, India, and the Americas. This is a full-stack role with primary focus on infrastructure. You will spend roughly two-thirds of your time on telemetry infrastructure, data pipelines, Kubernetes, and DevOps work, with the remainder on services and product surfaces. You will design and own distributed systems end-to-end, from event ingestion through production operation, and will carry the pager for systems you build. Key responsibilities include: **Telemetry Infrastructure & Data Pipelines:** Design, build, and own backend services for distributed data streaming and processing handling high-volume, high-cardinality event data with predictable latency and no silent data loss. Build and maintain Kafka-based streaming pipelines feeding ClickHouse. Own data modeling and query performance in ClickHouse, including partitioning, sort keys, materialized views, retention, and cost optimization. Partner with product and platform teams to shape telemetry requirements and drive solutions to production. **Kubernetes, DevOps & Production Ownership:** Deploy, scale, and operate services in production Kubernetes environments using Helm-based deployments and CI/CD pipelines. Own observability for your systems—meaningful metrics, useful logs, real tracing, and alerts that fire on customer impact. Debug and resolve production incidents independently, participate in on-call rotations, run root-cause analysis, and operate against defined service level objectives. **Full-Stack Engineering & Technical Leadership:** Own the complete development lifecycle for your work, building backend services and APIs that expose telemetry to internal consumers, product surfaces, and AI agents. Contribute to analytics frontends—dashboards, funnels, exploration tools—with attention to performance against large result sets. Improve web and mobile capture SDKs. Lead design of complex, multi-service features across team boundaries, drive decisions to closure, and write design docs and postmortems. Raise the bar through code review, test strategy, automation coverage, and mentoring. The role emphasizes self-direction and technical depth. You will own technical decisions within your domain and escalate only genuine ambiguities. Incident response management is a critical skill at Expert level; other competencies are at Experienced level. **Requirements:** **Mandatory Minimums:** - 8+ years of backend software engineering experience - 4+ years designing and operating distributed systems - 3+ years Kubernetes in production, including Helm and CI/CD - 3+ years Kafka or equivalent streaming and data pipelines - 3+ years on-call and production ownership - Bachelor's degree in computer science, software engineering, or closely related technical field - (A master's degree may offset up to one year of experience minimum; equivalent practical experience considered where technical depth is demonstrable) **Required Qualifications:** - Strong backend development experience in Java, Python, Go, or equivalent, used in production at scale - Hands-on experience building distributed systems for data streaming, processing, and storage - Production experience deploying and managing services on Kubernetes, including Helm and CI/CD pipelines - Working knowledge of Kafka or equivalent messaging and streaming system - Solid SQL skills and hands-on experience with a columnar or analytical data store, including query optimization and physical data modeling - Demonstrated depth in incident response and customer escalations - Sufficient front-end capability to build and debug data-heavy interfaces: JavaScript or TypeScript with React or Angular **Desired Qualifications:** - Production experience with ClickHouse or another OLAP or columnar store (sharding, replication, materialized views, cost tuning) - Exposure to observability tooling: OpenTelemetry, metrics and logging stacks, alerting platforms - Experience building product analytics or telemetry platforms, or with stream processing frameworks such as Flink or Spark - Client-side instrumentation: web or mobile SDK development, event capture, session and consent handling **AI Skills (Experienced proficiency required):** - Confident use of AI coding assistants such as GitHub Copilot or Cursor, with critical eye to review AI-generated code for logic errors, security anti-patterns, and missed edge cases - AI-assisted debugging and trace analysis, with clear understanding of data privacy rules for prompts - Integrating LLM APIs into production features; knowing when a model is the wrong tool. Working knowledge of RAG pipelines, embedding stores, and agent frameworks; familiarity with Model Context Protocol is a plus - Exposing telemetry to AI agents through safe tool interfaces: predictable schemas, bounded results, clear error semantics, cost controls - Designing safeguards for non-deterministic components: retry logic, circuit breakers, human-in-the-loop checkpoints; hardening against latency spikes and hallucinated output - Instrumenting AI features in production and judging quality from telemetry and evaluation

Similar roles