SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mistral is seeking an AI Scientist to join the AI4Engineering Science team, working on agentic systems that can reason about and execute real engineering tasks. The role focuses on building pre- and post-training data for Mistral's frontier LLMs, designing verifiers and evaluation criteria, and developing agent architectures for multi-step engineering workflows.
Key responsibilities include designing pretraining, SFT, and RL data for engineering tasks across CAD, CAE, and semiconductor/EDA domains; defining verifiers that capture what "correct" means for each task; designing and improving agent architectures for tool use, planning, and error recovery; building evaluation benchmarks and diagnostic tooling to identify model failures; and collaborating with domain experts to identify high-value engineering workflows.
The ideal candidate has deep, hands-on machine learning expertise with particular strength in LLM development. You should have demonstrated experience running, debugging, and validating real engineering workflows in at least one of CAD, CAE, semiconductor simulation, or EDA. Strong Python coding skills and comfort with Linux/HPC environments are essential. You are self-directed, collaborative, and able to communicate technical ML and engineering concepts to diverse audiences.
Desirable qualifications include experience building or fine-tuning agentic systems (tool use, multi-step planning, agent orchestration); experience with reward modeling and preference-based training methods (RLHF/RLAIF/RLVR); industrial or academic experience with specific CAE tools (SolidWorks, CATIA, Fluent, Abaqus, LS-DYNA, STAR-CCM+) or EDA tools (Cadence, Synopsys, Siemens); contributions to large open-source or industry codebases; and publications in engineering and ML venues (NeurIPS, ICLR, JFM, AIAA).
This is early-stage, foundational work where you'll shape the data, verifiers, and agent scaffolding that determine whether these systems become reliable. You'll work closely with domain experts and the broader research team to translate domain knowledge into training signals and evaluation benchmarks that measure genuine task competence.