SlipstreamJobsFresh Startup & VC-Backed Jobs

AI Engineer, Product

Mistral - Paris, Île-de-France, France - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Mistral is a full-stack AI company providing frontier models, developer tools, applications, and compute infrastructure. They partner with enterprises across finance, manufacturing, defense, healthcare, and the public sector to co-create customized AI systems. As an AI Engineer embedded in a product team (search, chat, documents, or audio), you will own AI quality end-to-end for your domain. Your responsibilities include designing and running rigorous evaluations tailored to your product area—reference tests, heuristics, and model-graded checks. You'll define and track metrics that matter: task success, helpfulness, hallucination proxies, safety flags, latency, and cost. You will own prompt and orchestration design as a core part of your work, writing and iterating on prompts and system prompts. You'll run A/B tests on prompts, models, and configurations, analyze results, and make data-driven rollout or rollback decisions. Setting up observability for LLM calls—structured logging, tracing, dashboards, and alerts—falls within your scope. Operating model releases is a key responsibility: managing canary and shadow traffic, sign-offs, SLO-based rollback criteria, and regression detection. You'll improve core behaviors in your product area, whether memory policies, intent classification, routing, tool-call reliability, or retrieval quality. You'll create templates and documentation so other teams can author evaluations and ship safely, and partner with the Science team to diagnose regressions and lead post-mortems. You should have 3–4 years of experience, with backgrounds including ML engineers moving closer to product or software engineers with real AI/ML production experience. Strong TypeScript or Python skills are required. You need production LLM experience with prompts, tool/function calling, and system prompts. Hands-on experience with evaluations and A/B testing is essential—you should be able to design metrics, not just run them. You must be comfortable implementing directly in product code, not only notebooks, and have observability experience with logging, tracing, dashboards, and alerting. A product mindset—forming hypotheses, running experiments, interpreting results, and shipping—is critical. Clear communication, autonomy, and orientation toward production impact are expected. Ideal candidates also have safety systems experience (moderation, PII handling, guardrails), release operations expertise (canary/shadowing, automated rollbacks, experiment platforms), or prior work on search ranking, chat systems, document AI, or audio ML features.

Similar roles