SlipstreamJobsFresh Startup & VC-Backed Jobs

Forward Deployed Engineer (Training)

Baseten - San Francisco, CA, USA - In-office - posted 2026-08-19

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Baseten powers mission-critical inference for leading AI companies including Cursor, Notion, Abridge, and Writer. The company recently raised $1.5B in Series F funding and is building the platform engineers rely on to ship AI products at scale. As a Forward Deployed Engineer, you will work directly with the world's largest and fastest-growing AI companies, owning their technical outcomes on Baseten and tackling the hardest problems in serving and improving models at scale. You'll act as each account's de facto CTO, taking customer objectives from vague concepts through to shipped production systems. Key responsibilities include: framing problems and defining specifications with customers; building and validating proofs of concept; designing evals and benchmarks to identify performance gaps, then closing those gaps through inference optimization, post-training improvements, or eval refinement; serving as first responder to mission-critical failures with full accountability; building internal systems and tooling to accelerate future engagements; shaping the product roadmap based on customer needs and shipping features directly into Baseten's codebase; and managing multiple concurrent accounts while keeping stakeholders aligned. You'll need minimum 1-2 years of software engineering experience shipping and maintaining code in large production systems, with breadth across the stack. Strong debugging skills for complex production issues, comfort with ambiguous technical problems, and clear communication on complex topics are essential. You should have genuine curiosity about AI inference and training infrastructure, and willingness to participate in on-call rotations. Ideal candidates may have depth in infrastructure domains (storage, networking, cloud layers), experience operating distributed compute platforms (Kubernetes, Slurm, Ray) for GPU workloads, detailed understanding of LLM architectures and modern inference engines (vLLM, TensorRT-LLM, SGLang), ability to profile and optimize GPU workloads, hands-on post-training experience (SFT, RL), or operational depth in incident response and distributed systems debugging.

Similar roles