SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Snowflake's ML Platform team is building Cortex Training, an LLM post-training platform that transforms GPU capacity into a simple, composable service for customers to adapt open-weight foundation models to their business needs. The platform handles distributed-systems complexity including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput optimization.
You will design and build across the full stack—from public training APIs and SDK through the control plane to the GPU data plane. You'll scale distributed systems that make GPU compute serverless, including multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools with built-in fault tolerance. You'll drive end-to-end performance at scale, keeping training, inference, and RL loops fast while maintaining GPU saturation under heavy concurrent load. You'll also productionize research building blocks by partnering with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components for enterprise-scale deployment.
The team ships fast while maintaining high reliability standards and works alongside the researchers behind DeepSpeed. This role requires someone who thrives in the ML infrastructure layer with solid understanding of LLMs and post-training workflows.
REQUIREMENTS:
- 6+ years building and shipping production ML systems
- Strong distributed systems and infrastructure foundation—designing scalable, fault-tolerant services and operating them on Kubernetes in production
- Familiarity with GPU and LLM infrastructure (e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM); ability to debug across data, infrastructure, and GPU layers
- Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency
- BS in Computer Science or related field (MS/PhD a plus)
- Bonus: hands-on LLM post-training/modeling experience