SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Baseten is a Series F AI infrastructure company ($1.5B valuation) that powers mission-critical inference for leading AI companies like Cursor, Notion, Abridge, and Writer. The company combines applied AI research, flexible infrastructure, and developer tooling to help frontier AI companies bring cutting-edge models into production.
You'll join the Training Product team as an AI Engineer, building agentic product features for customers training and post-training frontier models on Baseten's platform. This is a hands-on role with significant autonomy where you'll identify problems worth solving, design and build the harnesses, execution flows, and guardrails that make AI systems reliable in production, and own the results.
Key responsibilities include: shipping agentic product experiences (chat, assistant interfaces) from prototype to GA; designing reliable AI systems for production; building internal automation and tooling that increases team velocity; partnering with research engineers to turn internal research workflows into customer-facing features; defining and instrumenting evals to measure output quality improvements; implementing features end-to-end across the stack (API, backend, agent orchestration, frontend); using Baseten's own products to develop customer workflow intuition; identifying opportunities to replace manual processes with AI; and resolving customer issues with urgency.
You'll need 5+ years shipping software applications with demonstrated experience building AI/LLM-powered products, agents, or agentic workflows that real users depend on. Strong software engineering fundamentals, ability to build accurate mental models of systems, Python proficiency plus fluency in another language, comfort working autonomously in fast-moving environments, and strong communication bridging technical depth and business needs are essential.
Nice-to-haves include founding/early-stage startup experience, experience with evals and agent observability, familiarity with agent frameworks (LangChain, Claude Code, MCP), knowledge of model development methods (fine-tuning, RL, synthetic data, LoRA), and frontend fluency.