SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mistral is building full-stack AI solutions, from frontier models to developer tools and compute infrastructure. This Research Engineer role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference across Mistral's distributed infrastructure.
You will develop the infrastructure enabling researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions. Responsibilities span the full ML lifecycle: workload scheduling, capacity management, platform APIs, observability, and production operations. You'll take ownership of critical systems and transform complex infrastructure into reliable, self-service capabilities.
Key responsibilities include:
- Building services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference
- Orchestrating GPU workloads with queueing, admission control, quotas, priorities, preemption, and topology-aware placement
- Managing heterogeneous GPU resource provisioning, allocation, and utilization across clusters
- Enabling multi-cluster execution with intelligent workload placement based on capacity, data locality, and hardware requirements
- Creating self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce
- Optimizing GPU utilization, scheduling latency, workload startup time, and infrastructure efficiency
- Building observability, failure recovery, and operational tooling for critical ML workloads
- Participating in on-call rotations and troubleshooting across applications, schedulers, networking, storage, and GPU infrastructure
Required qualifications: 4+ years in ML infrastructure, distributed systems, or Kubernetes platform engineering. Proficiency in Python or Go with production-grade distributed systems experience. Strong Kubernetes knowledge including controllers, operators, CRDs, scheduling, networking, and storage. Familiarity with tools like Kueue, Karpenter, Volcano, and Kyverno. Understanding of distributed ML workloads (training, fine-tuning, evaluation, checkpointing, batch inference). Experience with GPU infrastructure, PyTorch, CUDA, NCCL, and high-performance networking. Knowledge of quotas, priorities, preemption, gang scheduling, and workload admission. Ability to diagnose performance and reliability issues across the full stack. Strong focus on developer experience and ability to thrive in fast-moving, ambiguous environments shaped by frontier AI research.