SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 170,000 - 205,000 / annual
Crusoe is a vertically integrated AI infrastructure company building the next generation of managed AI services. The company owns and operates each layer of the stack—from electrons to tokens—to power the world's most ambitious AI workloads, with an energy-first approach that optimizes both performance and sustainability.
You will join the Crusoe Cloud Managed AI team as a Senior Software Engineer with a pivotal role in shaping the architecture and scalability of next-generation managed AI services. This is a hands-on technical leadership position where you'll lead the design and implementation of core systems including resilient fault-tolerant queues, model catalogs, and scheduling mechanisms optimized for cost and performance.
Key responsibilities include:
- Design and implement core AI services including fault-tolerant queues for task distribution, model catalogs for managing and versioning AI models, and cost/performance-optimized scheduling mechanisms
- Architect and scale infrastructure to handle millions of API requests per second across thousands of customers
- Develop deep expertise in Kubernetes or other distributed cluster orchestration systems
- Implement robust monitoring and alerting to ensure system health and 24/7 availability
- Collaborate closely with product management, business strategy, and other engineering teams to define the AI platform roadmap
- Influence long-term vision and architectural decisions of the platform
- Contribute to open-source AI frameworks and actively participate in the AI community
- Prototype and rapidly iterate on emerging technologies and new features
This role is based in San Francisco or Sunnyvale, CA, with required in-office presence.
REQUIREMENTS:
- Advanced degree in Computer Science or Engineering
- 4-5+ years of industry experience with demonstrated history of consistent success leading a varied portfolio of initiatives
- Strong experience with distributed systems, cloud services (compute, storage, networking, database), and delivering early-stage projects quickly
- Experience with Generative AI (LLMs, Multimodal) and AI infrastructure (training, inference, ETL pipelines)
- Proficiency with container runtimes (e.g., Kubernetes), microservices, REST APIs, gRPC, and full software development lifecycle including CI/CD
PREFERRED QUALIFICATIONS:
- Proficiency in Golang, Python, or Rust for production services
- Contributions to open-source AI projects (e.g., VLLM)
- Experience with performance optimizations on GPU systems and inference frameworks