SlipstreamJobsFresh Startup & VC-Backed Jobs

Head of Engineering

Inferact - San Francisco, CA, United States - In-office - posted 2026-08-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Inferact, founded by the creators and core maintainers of vLLM, is building the world's AI inference engine. The company sits at the intersection of models and hardware, focused on making inference cheaper and faster. As Head of Engineering, you will build and lead the engineering organization developing the systems that power vLLM and Inferact's products. This is a leadership role requiring genuine technical credibility at the inference layer—deep understanding of GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design. Key responsibilities include: - Recruiting and developing rare ML systems talent, particularly senior engineers, staff-level individual contributors, and PhD-level researchers - Partnering with founders to scale a senior-heavy, highly specialized engineering team while maintaining technical rigor, speed, and ownership - Translating ambitious research and infrastructure work into focused execution plans - Strengthening team operations and engineering practices - Delivering reliable, high-performance inference across models, hardware, and deployment environments - Remaining technically engaged to identify risks, pattern-match on difficult problems, and unblock teams without becoming a bottleneck - Representing the engineering organization credibly with open-source contributors, hardware partners, cloud providers, customers, and investors Required qualifications: - Bachelor's degree or equivalent in computer science, engineering, machine learning, systems, or related field - Engineering leadership experience building and scaling specialized teams in LLM inference, ML systems, GPU/accelerator software, or distributed systems - Deep technical credibility including hands-on understanding of inference runtimes, GPU optimization, kernels, memory/communication bottlenecks, and hardware-software tradeoffs - Strong record recruiting and retaining senior engineers, staff-level ICs, and research-adjacent talent - Ability to translate technically ambitious work into clear priorities, ownership, and execution plans - Experience remaining close to technical work to identify risks and unblock teams Preferred experience includes leading teams on LLM serving, vLLM, SGLang, GPU kernels, compiler/runtime systems, or distributed AI infrastructure; scaling senior-heavy organizations in narrow talent markets; integrating research/PhD talent into production teams; and contributing to open-source ML systems projects.

Similar roles