SlipstreamJobsFresh Startup & VC-Backed Jobs

Director, AI Systems Solutions Engineering

Tensordyne - Sunnyvale, CA, United States - In-office - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of generative AI workloads. The platform combines purpose-built silicon, optimized networking, and memory architecture into a tightly integrated system for large-scale AI inference, serving hyperscalers, neoclouds, frontier model developers, and enterprises. As Tensordyne moves from system development into silicon bring-up, customer validation, beta deployments, and production rollout, the company is building a technical customer organization that will sit between engineering teams and deploying customers. You will own and grow Tensordyne's most important technical customer engagements. This is a senior, highly technical role for someone who understands modern AI infrastructure from model architecture through accelerator hardware, distributed inference, serving software, and datacenter deployment. You will work directly with frontier model builders, hyperscalers, neoclouds, developers, infrastructure partners, and strategic customers as they evaluate and deploy Tensordyne systems. Key responsibilities include: - Own strategic technical customer engagements from initial architecture discussions through benchmarking, evaluation, integration, deployment, and expansion - Define how Tensordyne demonstrates system performance across KPIs like throughput, tokens/sec/user, TTFT, memory utilization, power efficiency, system density, and inference metrics - Work with customers and model developers to understand HW and model architectures, serving requirements, context lengths, parallelism strategies, and inference optimization requirements - Maintain technically rigorous understanding of Tensordyne performance relative to leading GPU and AI accelerator platforms; ensure customer-facing comparisons are credible and reproducible - Partner with compiler, runtime, kernel, systems, and SDK teams to bring customer models onto the Tensordyne platform and identify performance improvement opportunities - Lead technical PoCs, remote evaluations, on-premises beta deployments, integration programs, and production readiness efforts - Translate recurring customer requirements into clear priorities for SDK, compiler, runtime, inference server, model support, orchestration, networking, and system architecture teams - Stay deeply current on model architectures, inference techniques, accelerator roadmaps, serving frameworks, competitive systems, and AI infrastructure market changes - Build and lead a small team of elite Sales and Solutions Engineers capable of independently managing sophisticated technical engagements with demanding AI infrastructure customers - Turn early customer engagements into repeatable benchmarks, evaluation frameworks, reference architectures, deployment playbooks, documentation, and technical collateral This is not a traditional pre-sales engineering role. The team will operate at the frontier of a rapidly changing technology landscape, working with constantly evolving new models and requirements. You will be equally comfortable in customer architecture reviews, helping prioritize product capabilities and roadmaps, and leading technical evaluations with hyperscalers and frontier AI companies. REQUIREMENTS: - Deep understanding of modern AI inference systems, including LLM and multimodal architectures - Strong knowledge of AI accelerator and system architecture, including compute, memory hierarchy, interconnect, parallelism, and distributed inference - Experience reasoning about inference performance across latency, throughput, memory bandwidth, utilization, batching, context length, prefill, decode, and system scaling - Hands-on familiarity with modern AI frameworks and serving environments such as PyTorch, vLLM, SGLang, Triton, or comparable systems - Experience working across the boundary between AI software and accelerator hardware, ideally including GPUs, custom silicon, or emerging AI accelerators - Experience benchmarking and optimizing workloads on large-scale AI infrastructure - Strong understanding of production inference techniques including quantization, tensor/model/expert parallelism, disaggregated serving, KV-cache management, distributed execution, and related optimization strategies - Demonstrated ability to engage technically sophisticated external organizations including senior technical leaders at hyperscalers, cloud providers, model developers, AI infrastructure companies, or large enterprise engineering teams - Experience leading high-performing Solutions Engineering, Field Engineering, Forward Deployed Engineering, or comparable technical customer teams - Ability to operate effectively in a fast-moving environment where product, software stack, competitive landscape, and customer requirements are evolving simultaneously Particularly relevant experience includes backgrounds from organizations building or deploying GPU or custom AI accelerator platforms, large-scale AI inference infrastructure, frontier or foundation models, hyperscale cloud infrastructure, AI serving and orchestration platforms, compiler/runtime/distributed AI systems, or high-performance computing. Experience bringing a new accelerator architecture or AI infrastructure platform from early access through customer validation and production deployment is especially valuable.

Similar roles