SlipstreamJobsFresh Startup & VC-Backed Jobs

Performance Engineer, Inference

Sarvam AI - Bengaluru, India - Hybrid - posted 2026-08-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sarvam AI is hiring a Performance Engineer to own the production serving path for large distributed language models end-to-end. You will be part of the Performance Engineering team, which owns the critical metrics the company plans against: serving latency, throughput, cost per token, and GPU utilization across a multi-node, multi-tenant fleet of NVIDIA Hoppers and Blackwells. In this role, you will be source-level fluent in at least one major inference serving framework (SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM) and able to read and modify these runtimes where stock behavior doesn't fit Sarvam's workloads. You will operate and extend a distributed-serving stack at depth, including disaggregated prefill-decode across nodes, distributed KV/cache transfer, and cross-node routing and scheduling. You will integrate artifacts from the model and kernel teams into a running multi-node, multi-tenant stack and build and train your own speculators—draft models, distillation from target models, and acceptance-rate tuning against live serving distributions—rather than only wiring in stock implementations. Your scoreboard includes TTFT (p50/p95/p99), TPOT, throughput, GPU utilization, and cost per million tokens. You will produce and defend the latency and throughput numbers the company plans against, and spend significant time in cross-team work with architecture, the kernels team, the model team, and SRE. Required qualifications: 5+ years in ML systems with 2+ years on inference serving at production scale, with concrete outcomes (tokens per day, throughput wins, p99 reductions). You must have production experience serving 100B+ parameter models across multi-node tensor, pipeline, or expert parallelism. You need source-level fluency in one of the four major serving frameworks and reading-level familiarity with the others. You should have operated and extended a disaggregated prefill-decode stack, distributed KV/cache transfer, and cross-node routing in production. Speculative decoding experience is essential—you should have trained draft models or speculators, distilled them from target models, measured acceptance rates against real serving distributions, and composed speculation with the rest of the stack. Deep understanding of KV cache internals (block tables, copy-on-write, prefix sharing, fragmentation), TP/PP/EP, NCCL primitives, multi-tenant serving (model co-location, MIG/MPS isolation), and profiling tools (Nsight Systems, framework tracing, py-spy/perf) are required. C++ and CUDA read-and-modify competency and on-call ownership of an inference SLO are expected. Strong pluses include upstream contributions to major serving frameworks, direct production experience with Dynamo or llm-d at scale, operating a forked runtime in production, published or shipped speculator work, and experience with MoE serving, long-context models (128K+), multi-model serving, or Indic/multilingual workloads.

Similar roles