SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 300,000 - 350,000 / annual
Bespoke Labs is an applied AI research lab pioneering data and reinforcement learning environment curation for training and evaluating agents. The company has recently curated Open Thoughts, one of the best open reasoning datasets used by frontier labs, trained state-of-the-art specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to perform multi-turn tool-calling with reinforcement learning.
This is a delivery-heavy role focused on standing up the enterprise post-training capability. Demand is inbound, and the goal is to ship custom, high-performing models to 2–3 paying enterprise customers by year-end. Over the first 6–12 months, you will execute the work required to solve real-world enterprise problems while helping build the scalable product underneath. You will not be running abstract experiments or training models solely for benchmarks. Instead, you will sit directly at the intersection of enterprise demand and applied post-training.
Key responsibilities include:
- Ship enterprise-grade models by post-training, fine-tuning, and aligning open-weight and proprietary base models for complex business domains, ensuring reliable production performance.
- Build bespoke evaluation suites by defining what "quality" means for subjective domain-specific tasks, creating custom benchmarks, and calibrating LLM judges against human domain experts.
- Curate and filter high-impact datasets by building and running production data flywheels combining real production traces, human labeling, and synthetic augmentation with strict filtering standards.
- Manage and mitigate regression risk by rigorously tracking downstream performance to ensure fine-tuning for new behaviors doesn't degrade core capabilities or reasoning.
- Engage directly with stakeholders including product teams and enterprise customers to understand requirements, translate domain preferences into technical eval metrics, and explain model behavior.
- Deploy for cost and privacy by fine-tuning open-weight architectures (Llama, Qwen, Mistral, DeepSeek) to hit strict enterprise latency, cost, and privacy targets.
- Direct frontier tools and workflows by leveraging state-of-the-art post-training techniques, preference tuning, and data curation tooling to maximize output quality and delivery speed.
You will be measured on production impact, robustness, and delivery of models shipped to real users, not on published papers.
REQUIREMENTS:
- A record of shipped models: you have post-trained at least one LLM that was deployed to real users in production, and can demonstrate how you measured its success.
- End-to-end eval ownership: demonstrated experience building benchmarks, creating eval datasets, and getting stakeholders to agree on clear metrics for complex or subjective tasks.
- Deep understanding of regression risks: ability to explain how training a model on new behaviors impacts existing capabilities and how to prevent degradation.
- Customer-facing or product empathy: experience collaborating directly with non-ML stakeholders, enterprise customers, or product managers to turn requirements into model behavior.
- Strong software and ML fundamentals: fluency in modern post-training frameworks, data processing pipelines, and code bases designed for production deployment.
- Ownership mindset: complete responsibility for the full post-training lifecycle—from raw data to model deployment and failure analysis—without requiring close supervision.
Additional experience that strengthens candidacy:
- Post-training conversational or task-oriented assistants (support agents, multi-turn chat, tool-using agents).
- Building LLM judges or reward models and calibrating them against human raters.
- Operating a production data flywheel: traces → labeling → synthetic augmentation → retrain.
- Extensive hands-on experience with open-weight models for cost, latency, or privacy optimization.
- Forward-deployed engineer, founder, or early-stage startup backgrounds.