SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Crusoe is a vertically integrated AI infrastructure company building the complete stack from energy to tokens. As a Staff Applied AI Inference Engineer, you will own the inference stack end-to-end, making large language models run faster, cheaper, and more reliably in production.
You'll spend your time on applied systems and performance work: profiling where time and cost go, bringing modern optimization techniques into real deployments, and diving deep into serving code when defaults aren't sufficient. The role balances hands-on engineering with customer-facing technical solutions work.
Key responsibilities include:
- Bringing current inference techniques into production and refining them across diverse models and traffic patterns
- Designing and optimizing serving architectures, including prefill/decode disaggregation and request routing
- Working down the serving stack from frameworks like vLLM and SGLang to CUDA kernels, profiling and analyzing performance bottlenecks
- Adapting optimization methods across many ML models with emphasis on large language models
- Profiling and tuning deployments against clear latency, throughput, and cost targets
- Partnering with customer engineering teams to tailor deployments and move workloads from proof of concept to production
- Building software and product features around the inference stack using Python and general-purpose languages
- Owning delivery end-to-end from experiment through production optimization
- Working through ambiguity, making sound technical tradeoffs, and avoiding unnecessary complexity
Required qualifications: Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field. Hands-on production experience with Python or C++ (Python preferred). Familiarity with LLM optimization for high throughput/low latency inference. Comfort with vLLM or SGLang. Strong understanding of GPU architecture and behavior. Clear interest and hands-on experience with large language models. Working knowledge of AI/ML pipelines and the full ML development and deployment path. Strong communication skills, especially explaining technical topics to customers and teammates.
Bonus: Track record optimizing software systems for LLMs, CUDA experience, strong software engineering fundamentals with AI/ML inference systems shipping record, Docker/Kubernetes experience, prior customer-facing AI/ML project work.