SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sarvam AI is hiring a Performance Engineer to own the kernel layer of its serving infrastructure. You will author custom CUDA, DSL-based, and PTX kernels that optimize performance beyond what stock libraries (cuBLAS, cuDNN, FlashAttention, Triton) provide. The Performance Engineering team owns the serving speed, cost efficiency, and GPU utilization metrics that the rest of the company plans against, working across a multi-node, multi-tenant fleet running H100/H200/B200 GPUs serving multiple model families including small and large LLMs, Mixture-of-Experts, streaming Indic ASR, and multimodal systems.
You will be responsible for closing performance gaps by authoring kernels that beat published baselines on real production workloads. When your code ships, production p99 latency moves measurably, and you own explaining why. This is a deep, high-leverage role requiring genuine kernel-authoring expertise, not just kernel usage.
Required qualifications: 5+ years in ML systems with 2+ years authoring production CUDA kernels. You must have shipped at least one kernel in production that beat the prior baseline by a measurable margin. You need CUDA expertise at the kernel-authoring level (thread-block sizing, shared-memory layout, warp primitives, async copies like cp.async and TMA, MMA selection). You should be comfortable modifying and extending CUTLASS/CuTe DSL, including the layout algebra. PTX proficiency at debug-and-modify level is required—you have inserted hand-written PTX where the compiler fell short. You must be fluent with Nsight Compute and Systems tools, able to read roofline plots and propose fixes. You have authored or significantly modified at least one attention kernel (FlashAttention-family, paged, MLA, sliding-window, or sparse). Multi-architecture awareness is essential: you understand what changes from Hopper to Blackwell (TMA, WGMMA, tcgen05).
Strong pluses include communication kernel expertise (NCCL/NVSHMEM authoring, custom collectives, expert-parallel dispatch, AFD bipartite comms, KV transfer), open-source kernel contributions (FlashAttention, CUTLASS, vLLM/SGLang, DeepEP, Mooncake, or Triton), advanced features like tcgen05, TMA, CTA-cluster launch, distributed shared memory, async pipelining, and FP4/microscaling paths, plus Grace-side host-path optimization on GH200/GB200.