SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
vCluster Labs is hiring a Senior Inference Engineer to own the inference layer of their AI infrastructure platform from the ground up. You will be the first engineer dedicated to inference, partnering directly with the CTO to build production-grade, query-to-response pipelines running at scale on GPU infrastructure.
Key responsibilities include:
- Deploying LLMs to production across GPU infrastructure, owning the full pipeline from customer query to served response
- Standing up and operating serving infrastructure using frameworks like vLLM, SGLang, or TensorRT-LLM
- Optimizing for scale through quantization, batching, caching, and routing to manage latency and cost
- Building real infrastructure in Python or Golang (engineering-focused, not research or data-science)
- Owning the inference platform roadmap alongside the CTO and Product team, shaping the direction as the space evolves
About vCluster Labs: A venture-backed startup (raised $28M+ from Khosla Ventures and others) providing the #1 platform for AI infrastructure. The company powers 100,000+ GPUs and 1M+ CPUs across 50+ AI clouds and Fortune 500 companies. Headquarters in San Francisco with a distributed, remote-first team of 40+ infrastructure engineers. The company is behind vCluster, an open-source Kubernetes tenant isolation technology with 11,000+ GitHub stars.
Requirements:
- Production LLM serving experience: deployed and served LLMs using vLLM, SGLang, or TensorRT-LLM, ideally at a company built around inference at scale
- Inference optimization hands-on experience: quantization, batching, caching, and routing (not just familiarity)
- Strong engineering skills in Python or Golang with real production code experience
- Strong communication skills, explaining technical concepts to engineers and non-technical stakeholders
Bonus qualifications:
- Familiarity with containerized environments (Docker, Kubernetes)
- Hands-on generative AI experience with common ML frameworks (PyTorch, Transformers)
- Understanding of the GPU stack: CUDA, NCCL, drivers, and related libraries
- Knowledge of model architectures and fine-tuning approaches
- Experience with NVIDIA Dynamo