SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Tensordyne is an AI systems company building high-performance, low-power generative AI inference systems through custom silicon, hardware, and software. The company enables multimodal generative AI inference acceleration at scale for hyperscaler and neocloud data center customers.
As a Software Engineer on the Low-Level Runtime team, you will develop and maintain components of an ML model execution framework optimized for Tensordyne's custom hardware. This is a hands-on role focused on building low-level assembler architecture and porting ML models for execution on the company's specialized inference platform.
Key responsibilities include:
- Maintain and develop the low-level ML model execution framework written in Rust
- Enhance and optimize components and complete models designed for Tensordyne hardware
- Drive performance optimization of ML models to maximize throughput and hardware utilization
- Debug and productize models across both virtual and physical hardware environments
Preferred qualifications include experience with Rust, familiarity with high-level ML concepts (LLMs, attention architectures), low-level systems programming (embedded systems), and experience with LLM engineering workflows. Nice-to-have skills include Google Cloud Platform experience, Python proficiency, and deep knowledge of modern attention mechanisms.
Tensordyne is well-funded with headquarters in Sunnyvale, CA and Munich, Germany, plus distributed teams across North America and Europe. The company emphasizes comprehensive benefits, competitive compensation, and flexible work arrangements.