SlipstreamJobsFresh Startup & VC-Backed Jobs

AI Systems Research and Development Engineer – LLM Inference Systems & Optimization

Snowflake - Bellevue, WA, United States - In-office - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Snowflake AI Research is seeking experienced systems engineers and researchers to advance the state of the art in LLM inference systems and optimization. This role focuses on building the next generation of high-performance inference systems that optimize both execution speed and efficiency, while enabling rapid adaptation to new models, architectures, hardware, and workloads. You will work across the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. The team explores cutting-edge techniques including adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization. Recent innovations include Arctic Inference (open-source), Shift Parallelism, SwiftKV, Arctic Speculator, SuffixDecoding, Jacobi Forcing, and Semi-Persistence technologies. Key responsibilities include designing and developing high-performance LLM inference systems; developing novel techniques to improve latency, throughput, memory efficiency, and cost; exploring advanced inference techniques; building adaptive systems that automatically optimize execution; applying AI-driven approaches to systems engineering; independently identifying and solving high-impact problems; designing distributed inference strategies across GPUs and nodes; developing efficient multi-model serving approaches; optimizing GPU kernels and operators; exploring model-system co-design; profiling and benchmarking end-to-end workloads; collaborating with model researchers and product teams; and publishing innovations through technical blogs and top-tier conferences. The ideal candidate brings 5+ years of experience in LLM inference systems, distributed AI systems, GPU systems, or high-performance computing. You should have strong understanding of modern LLM inference architectures, hands-on experience with frameworks like vLLM, SGLang, or TensorRT-LLM, experience designing inference runtimes, strong GPU architecture knowledge with CUDA/Triton experience, and familiarity with performance-oriented libraries. A Bachelor's in Computer Science or related field is required; Master's or PhD preferred. You'll collaborate with a world-class team including founding members of DeepSpeed, vLLM, and TensorFlow.

Similar roles