SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 250,000 - 300,000 / annual
Crusoe is a vertically integrated AI infrastructure company building the energy and compute stack for large-scale AI workloads. The company owns and operates every layer from power generation through serving, with a mission to accelerate AI abundance while solving the energy bottleneck.
In this role, you will own the inference stack end-to-end, making large language models run faster, cheaper, and more reliably in production. Your work is applied and hands-on: you'll profile performance bottlenecks, bring modern optimization techniques into real customer deployments, and dive deep into serving code when defaults fall short. You'll work with frameworks like vLLM and SGLang, optimize CUDA kernels, and tailor deployments to each customer's specific models, traffic patterns, latency targets, and cost constraints.
Key responsibilities include:
- Design and optimize serving architectures (prefill/decode disaggregation, request routing, etc.)
- Profile and analyze performance down to the kernel level, identifying and fixing bottlenecks
- Adapt optimization methods across ML models with emphasis on large language models
- Partner directly with customer engineering teams to move workloads from proof-of-concept to production
- Build software and product features around the inference stack using Python and general-purpose languages
- Own delivery end-to-end: from fuzzy goals through clear specs, fast experiments, and well-tested production results
- Work through ambiguity, make sound tradeoffs on tooling, and avoid unnecessary complexity
- Take pride in ownership and accountability, holding yourself and teammates to high standards
This is a hands-on engineering role with a customer-facing component and elements of product and technical solutions work. You will ship code that directly impacts real customer deployments.
REQUIREMENTS:
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field
- Hands-on production experience shipping code in Python or C++ (Python strongly preferred)
- Familiarity with LLM optimization methods for high-throughput, low-latency inference
- Comfort with modern LLM serving frameworks (vLLM, SGLang) and kernel-level profiling
- Firm grasp of GPU architecture and behavior
- Clear interest and hands-on experience with large language models
- Working knowledge of AI/ML pipelines and the full ML development/deployment path
- Strong communication skills, especially explaining technical topics to customers and teammates
BONUS:
- Track record of making software systems faster, especially for LLMs
- CUDA or comparable technology experience
- Strong software engineering fundamentals with record of shipping AI/ML inference systems
- Docker and Kubernetes experience
- Prior work building or tuning AI/ML projects in customer-facing settings