SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Tensordyne is an AI system solution company building high-performance, low-power generative AI inference systems. The company creates custom silicon, hardware, and software to enable multimodal generative AI inference acceleration at scale for hyperscaler and neocloud data center customers. Headquarters are in Sunnyvale, CA and Munich, Germany, with remote team members across North America and Europe.
As a Systems Performance Modeling Engineer, you will build models and tools that predict how generative AI inference workloads perform on Tensordyne systems, from single accelerators through rack, pod, and cluster scale. This is a hands-on role combining simulator development, experimental validation, and cross-stack performance analysis.
Key responsibilities include:
- Implement and extend simulation-based performance models for multimodal generative AI inference at rack, pod, and cluster scale, covering compute, memory, collective communication, and network fabric
- Model how serving strategies (tensor, pipeline, and expert parallelism, prefill/decode disaggregation, batching, KV-cache placement) interact with Tensordyne silicon and fabric topology, measuring effects on latency, throughput, and cost per token
- Build trace-capture and replay tooling that records real execution from the inference runtime and replays it under hypothetical silicon, system, and network configurations
- Model collective communication on multi-hop scale-out fabrics, including implementing custom collective algorithms for the topology
- Run calibration experiments on Tensordyne hardware, compare against model predictions, and debug sources of error
- Execute design-space sweeps and produce clear analyses for architects and engineering teams to inform ASIC, fabric, and system configuration decisions
- Generate performance projections supporting product and customer discussions
- Maintain the modeling codebase for speed, testability, and reproducibility
You will work closely with architects and silicon, hardware, networking, and software teams, moving across the stack from model graphs through collectives to network fabric to identify performance bottlenecks.
REQUIREMENTS:
- Hands-on experience building performance models, simulators, or analytical tools for ML workloads, distributed systems, or computer architecture
- Solid understanding of distributed ML execution, including parallelism strategies, collective communication (All-Reduce, All-Gather, All-to-All, etc.), and scaling behavior
- Working knowledge of system architecture across compute, memory, interconnect, and networking; ability to reason about bottlenecks between them
- Experience comparing model predictions against real measurements and debugging divergences
- Strong programming skills in C++ and Python, with emphasis on clean, testable, maintainable code
- Ability to take loosely defined performance questions, break them into experiments, and deliver results with minimal hand-holding
- Clear written and verbal communication, especially when presenting data and trade-offs to other engineers
- MS or higher in Computer Science, Computer Engineering, Electrical Engineering, or related field
NICE TO HAVE:
- Familiarity with LLM inference serving: batching, KV-cache management, disaggregated prefill/decode, latency/throughput trade-offs
- Experience modeling or benchmarking collective communication libraries (NCCL, RCCL, or similar) on real clusters
- Background in data center or HPC networking: topologies, RDMA/RoCE, congestion behavior
- Experience profiling ML workloads on accelerators (GPUs, TPUs, custom ASICs)
- Exposure to hardware/software co-design or early-stage architecture evaluation
- Publications or open-source contributions in ML systems, architecture, or networking