SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sail is building the foundational infrastructure for agentic AI systems. This role focuses on designing and implementing the distributed systems that make AI inference fast, reliable, and cost-efficient at global scale.
You will work on several core technical challenges: designing high-performance schedulers (admission control, queuing, priority, fairness, preemption, bin packing) that manage massive token queues across diverse hardware fleets; building global routing and traffic management systems with latency-aware dispatch, predictive autoscaling, and failover strategies; implementing LLM-specific routing optimizations such as KV caching strategies that trade memory for compute across GPU RAM, CPU RAM, and NVMe flash hierarchies; and building deep observability systems that trace every millisecond of system behavior to catch failures before customers notice them.
The ideal candidate has strong distributed systems fundamentals including concurrency, networking, databases, and performance engineering. You should be comfortable with the inherent complexity of distributed systems—understanding that correctness requires careful testing, edge-case analysis, and clear documentation. Experience with ML inference stacks (vLLM/SGLang), GPUs, or accelerators is a bonus.
Sail's interview process is designed to respect your time and assess fit thoroughly. You'll meet the CEO first, then the CTO for technical discussion, followed by an in-office interview day in San Francisco where you'll work on a realistic problem for 3-4 hours with immediate feedback, then present your approach and results. The company provides a modern office in downtown SF with excellent meals, Studio Displays, and a focus on removing friction from daily work.