SlipstreamJobsFresh Startup & VC-Backed Jobs

Inference Engineer

Hyperbolic - San Francisco, CA, USA - In-office - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Hyperbolic Labs is building an Open-Access AI Cloud that democratizes access to computing power through an innovative GPU marketplace and AI inference service. The company aims to make AI innovation universally accessible by optimizing idle computing resources globally. As an Inference Engineer, you will own the end-to-end design and deployment of inference capabilities on Forge, Hyperbolic's unified control plane. Your primary focus is enabling customers to consume model tokens without managing GPUs directly, while providing NeoCloud partners with a complete path to building their own token-factory offerings. Key responsibilities include: - Deploying and serving models across globally distributed clusters on heterogeneous hardware - Evaluating and selecting inference frameworks and serving engines for different workloads - Building production-ready deployment infrastructure including monitoring, gateways, and endpoints - Optimizing inference performance through autoscaling, KV-cache orchestration, and system tuning - Debugging customer inference issues and improving performance across the platform - Operating Kubernetes clusters in production at scale - Expanding scope into optimization, speculative decoding, and advanced inference techniques You will be the primary owner of inference at Hyperbolic, building the entire stack with significant influence over product direction and scope. This role requires strong foundational knowledge across the full inference stack—from request handling through token generation—rather than narrow specialization. Required qualifications include deep Kubernetes production experience, solid understanding of inference performance concepts (TTFT, disaggregated inference, speculative decoding, KV cache), familiarity with modern inference frameworks, working knowledge of NVIDIA Dynamo in distributed architectures, and proven ability to build products end-to-end from conception to serving real traffic. You should have strong self-initiative, comfort operating independently with minimal direction, and generalist instincts to take on adjacent work as needed. Preferred experience includes both inference deployment and optimization work, hands-on model optimization (quantization, batching, kernel tuning), understanding of RDMA and high-performance networking for distributed serving, experience with heterogeneous accelerators, direct customer support on inference debugging, and background at GPU cloud providers or AI infrastructure companies.

Similar roles