SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Crucibl is a profitable, fast-growing AI company building judgment-at-scale systems for enterprise decision-making. The $400B management consulting industry is being disrupted by AI that can handle the messy, ambiguous decisions driving 85% of the global economy. Crucibl raised a $10M seed from Tier 1 VCs and is scaling to meet a full client pipeline of Fortune 500 customers.
As a Member of Technical Staff - Research, you will push the frontier of how large language models reason through high-stakes business decisions and identify where they fail. You'll design novel methods for reasoning, evaluation, and calibration that go beyond standard benchmarks, translating open problems in reasoning, uncertainty, and multi-step decision-making into testable approaches.
You'll build evaluation frameworks that define and measure "good judgment" for models, designing experiments that reveal real failure modes rather than just leaderboard improvements. Your findings will directly inform product and applied team recommendations.
You'll partner closely with the founding team—bringing outside research thinking to a problem few labs focus on—and shape the technical vision and roadmap. You'll set the bar for what rigorous, applied research looks like in an AI-first organization as the team grows.
Must-haves: at least one undeniable signal of excellence (published research at a top venue, frontier lab role, or track record of novel technical contributions); deep understanding of how LLMs reason, fail, and can be evaluated; strong fundamentals in ML, statistics, or NLP; ability to move quickly from open-ended research questions to testable hypotheses.
Strong signals include: experience designing evaluation frameworks for reasoning or agentic systems; published or shipped work on uncertainty, calibration, multi-step reasoning, or LLM evaluation; experience in high-growth, high-ambiguity environments; owner mentality.
Not a fit: researchers needing clean, self-contained problems; pure benchmark-chasers; researchers without product instincts. You'll work with founders who bring Google-scale systems experience and deep domain expertise in high-stakes business decisions. Hybrid schedule: 3 days in-office with flexibility. Comprehensive salary and equity package, full health coverage, daily lunch and snacks.