SlipstreamJobsFresh Startup & VC-Backed Jobs

AI Researcher - Bolter

Improbable - Remote - Remote - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Bolter is an AI agent platform that enables users to describe work in plain language and have AI agents execute it end-to-end, creating shareable apps like trackers and dashboards. The company is at validation-sprint stage with a working product and is now scaling. As AI Researcher, you will make Bolter's agents trustworthy, capable, and useful through applied research with direct product impact. This role sits at the intersection of research and engineering, requiring you to design and run rigorous experiments, evaluate agent behavior, and translate findings into shipped features. Key responsibilities include: - Design and execute experiments testing agent reliability, context retention, multi-step task completion, and failure modes - Build and own the evaluation layer, creating evals that measure whether agents actually complete work correctly, not just generate plausible outputs - Research the frontier of LLM agents, tool use, and reliability, translating state-of-the-art insights into product decisions - Produce clear, actionable recommendations for the engineering team to ship - Prototype research findings into working products and hand off validated ideas - Shape the research roadmap as the product grows You'll work directly with the GM and product engineer in a lean team augmented by specialized AI agents. The core research challenge is ensuring agents that retain context, execute multi-step workflows, and create software can be reliably trusted to do so correctly. Ideal candidates have a track record in applied AI/ML research, particularly in LLM application design, agentic systems, evals, or production AI reliability. You should have strong experimental design skills, engineering fundamentals to prototype your own experiments, deep familiarity with LLMs and agentic failure modes, and excellent written communication. You're rigorous but pragmatic, comfortable with ambiguity, and can leverage AI agents as a force multiplier. Bonus experience includes LLM evals, hallucination mitigation, building agentic systems, and a public track record of shipped research.

Similar roles