SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mistral is building full-stack AI solutions spanning frontier models, developer tools, applications, and compute infrastructure. The company partners with enterprises across finance, manufacturing, defense, healthcare, and the public sector to co-create customized AI systems.
As a Research Engineer on the ML track, you will build and optimize the large-scale learning systems powering Mistral's open-weight models. You can choose between two paths: joining the Platform RE Team to enhance shared training frameworks, data pipelines, and cluster tooling used across all teams, or embedding within a research squad (Alignment, Pre-training, Multimodal, etc.) to translate cutting-edge research ideas into scalable, production-grade code.
Key responsibilities include accelerating researchers by owning heavy lifting in large-scale ML pipelines and building robust tools; bridging cutting-edge research with production by integrating checkpoints, streamlining evaluation, and exposing APIs; conducting experiments on latest deep-learning techniques including sparsified 70B+ model runs and distributed training across thousands of GPUs; designing, implementing, and benchmarking ML algorithms with clear, efficient Python code; and delivering prototypes that become production components for Le Chat and enterprise APIs.
You should have a Master's or PhD in Computer Science (or equivalent proven track record) with 4+ years working on large-scale ML codebases. Hands-on expertise with PyTorch, JAX, or TensorFlow is required, along with comfort in distributed training frameworks (DeepSpeed, FSDP, SLURM, K8s). Deep learning, NLP, or LLM experience is essential; CUDA or data-pipeline expertise is a bonus. Strong software-design instincts around testing, code review, and CI/CD are expected. The ideal candidate is a self-starter with low ego and strong collaborative skills.