SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Crusoe is a vertically integrated AI infrastructure company building the energy and compute foundation for large-scale AI workloads. As a Staff Applied AI Inference Engineer, you will own the inference stack end-to-end, optimizing large language models to run faster, cheaper, and more reliably in production.
Your core responsibilities include profiling and analyzing where time and cost are spent in inference deployments, implementing modern optimization techniques into real customer workloads, and diving deep into serving code and CUDA kernels when standard approaches fall short. This is hands-on systems and performance engineering work on some of the most demanding models in production today.
You will design and optimize serving architectures including prefill/decode disaggregation, request routing, and related approaches. You'll work across the full stack—from high-level frameworks like vLLM and SGLang down to GPU kernel-level optimization. Your work is applied and customer-centric: you'll partner directly with customer engineering teams to tailor deployments to their specific models, traffic patterns, latency targets, and cost constraints. You'll take workloads from proof-of-concept through to fully monitored production services, ensuring performance gains materialize in real deployments.
The role combines deep technical work with customer-facing elements and product collaboration. You'll experiment rapidly, shape fuzzy goals into clear specifications, run focused proofs of concept, and ship well-tested results. You'll own delivery end-to-end, from initial experiments through production optimization, using Python and/or C++ as primary languages.
Required qualifications include a Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field. You need hands-on production experience shipping code in general-purpose languages (Python preferred), strong familiarity with LLM serving frameworks and performance profiling down to kernel level, and deep understanding of GPU architecture and behavior. Clear interest and hands-on experience with large language models is essential, along with working knowledge of AI/ML pipelines and the full model development and deployment lifecycle. Strong communication skills—especially explaining complex technical topics to customers and teammates—are critical.
Bonus experience includes a track record optimizing software systems for speed (especially LLMs), CUDA or comparable GPU programming, strong software engineering fundamentals with shipped AI/ML inference systems, Docker and Kubernetes experience, and prior customer-facing AI/ML project work.