SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 200,000 - 400,000 / annual
Inferact is building vLLM as the world's AI inference engine, founded by the creators and core maintainers of vLLM. The company sits at the intersection of models and hardware, optimizing AI inference to make it cheaper and faster.
As a Forward Deployed Engineer, you will work directly with customers to deploy, integrate, debug, and optimize vLLM-powered inference systems across cloud, Kubernetes, GPU, networking, and model-serving environments. This is a hands-on engineering role focused on implementation, not pre-sales. You will move from architecture discussions into production deployment, own difficult customer problems end-to-end, and partner closely with core product and engineering teams to translate field learnings into reusable product capabilities. Your work directly impacts customer time-to-value and shapes how Inferact's platform evolves.
Key responsibilities include:
- Deploying and operating ML systems and model-serving infrastructure in production customer environments
- Debugging complex issues across application, runtime, infrastructure, networking, and distributed-system boundaries
- Working with sophisticated customer engineering teams to understand ambiguous technical environments and drive implementations to resolution
- Reasoning about latency, throughput, batching, model/runtime compatibility, scaling, reliability, and cost tradeoffs in production inference
- Identifying one-off customer problems versus systemic issues that should become reusable product capabilities
- Building strong technical communication and ownership across customer and internal teams
REQUIREMENTS:
Minimum qualifications:
- Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar
- Strong software engineering ability in Python, Go, TypeScript, or similar, with experience building production-quality integrations, tooling, services, automation, or prototypes
- Hands-on experience deploying or operating ML systems, model serving, AI infrastructure, cloud platforms, Kubernetes, or high-scale backend systems in production
- Ability to work directly with sophisticated customer engineering teams, understand ambiguous technical environments, and personally drive implementations and debugging to resolution
- Strong systems debugging skills across application, runtime, infrastructure, networking, identity, storage, observability, and distributed-system boundaries
- Ability to reason about latency, throughput, batching, model/runtime compatibility, scaling, reliability, and cost tradeoffs in production inference environments
- High ownership and strong technical communication, with judgment to distinguish one-off customer work from problems that should become reusable product capabilities
Preferred qualifications:
- Experience with vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, BentoML, or other LLM inference and model-serving systems
- Experience with NVIDIA or AMD GPUs, CUDA / ROCm, GPU scheduling, multi-GPU serving, or accelerator-backed infrastructure
- Experience deploying infrastructure software into enterprise, regulated, security-sensitive, or bring-your-own-cloud environments
- Experience building APIs, SDKs, CLIs, developer tooling, deployment platforms, control planes, or infrastructure products used by technical teams
- Experience profiling latency, throughput, concurrency, GPU utilization, bottlenecks, and performance regressions
- Contributions to open-source ML systems, inference infrastructure, cloud infrastructure, Kubernetes, or developer tooling
- Forward-deployed, customer engineering, field engineering, or highly technical solutions roles where you personally wrote and shipped code
- Built deployment playbooks, reference architectures, automation, or tooling that materially reduced customer time-to-production
- Resolved severe customer-facing production issues crossing multiple technical layers requiring close partnership with core engineering
- Turned repeated customer problems into reusable product features, abstractions, documentation, or platform improvements