SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Manager, Engineering - AI Inference

Crusoe - San Francisco, CA, United States - In-office - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 250,000 - 300,000 / annual

Crusoe is a vertically integrated AI infrastructure company building energy-efficient solutions for large-scale AI workloads. The company owns and operates the full stack from power generation through software, addressing the critical bottleneck of energy availability for AI compute. As Senior Manager of Engineering for AI Inference, you will lead an engineering team responsible for making large language models run faster, cheaper, and more reliably in production. This is a hybrid role combining people management with hands-on technical leadership. You will stay deeply involved in the inference stack end-to-end: profiling performance, implementing modern optimization techniques, and coding alongside your engineers. Key responsibilities include: - Leading team delivery from early experiments through production optimization, establishing performance goals and technical strategy - Designing and optimizing serving architectures, including prefill/decode disaggregation, request routing, and related approaches - Working across the serving stack from frameworks like vLLM and SGLang down to CUDA kernels, profiling to identify and fix performance bottlenecks - Adapting optimization methods across diverse ML models with emphasis on large language models - Profiling and tuning deployments against clear latency, throughput, and cost targets under real traffic - Partnering directly with customer engineering teams to tailor deployments to their specific models and constraints, moving workloads from proof-of-concept through production - Building software and product features around the inference stack using general-purpose languages (Python preferred) - Running fast experiments to validate approaches and shipping well-tested results quickly - Making sound technical tradeoffs and steering away from unnecessary complexity The role balances performance engineering with customer-facing technical solutions leadership, ensuring optimization gains deliver measurable business value. REQUIREMENTS: - 2+ years directly managing and leading engineering teams in high-performance or ML-focused environments - Strong hands-on software engineering experience with low-level optimization or ML infrastructure; continued desire to stay close to code and architecture - Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field - Production shipping experience with one or more general-purpose languages (Python or C++; Python strongly preferred) - Familiarity with LLM optimization techniques for high-throughput, low-latency inference - Hands-on experience with modern LLM serving frameworks (vLLM, SGLang) and kernel-level performance profiling and analysis - Solid understanding of GPU architecture and behavior - Clear interest and hands-on experience with large language models - Working knowledge of AI/ML pipelines and the full development-to-deployment path - Strong communication skills, especially explaining complex technical topics to customers and teammates BONUS: - Track record of optimizing software systems for speed, especially LLMs - CUDA or comparable technology experience - Strong software engineering fundamentals and record of building/shipping AI/ML inference systems - Docker and Kubernetes experience - Prior work building or tuning AI/ML projects in customer-facing settings

Similar roles