SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mistral is building full-stack AI solutions, from frontier models to developer tools and compute infrastructure. This role focuses on designing and operating next-generation data infrastructure to support massive-scale AI training and fine-tuning operations.
You will own the full lifecycle of critical infrastructure projects, including architecting the migration from legacy orchestrators to modern, cloud-native systems. Key responsibilities include:
• Building and scaling distributed compute and storage systems to support Mistral's growing training workloads
• Designing multi-cluster orchestration layers to optimize workload placement across diverse hardware and geographic regions
• Architecting the transition to modern storage formats (columnar standards) to handle fine-tuning datasets at exabyte scale
• Contributing to the internal training platform, ensuring seamless model training capabilities across Kubernetes and SLURM environments
• Implementing metadata and lineage systems to provide visibility across increasingly complex data and model pipelines
• Managing cloud-native deployments and ensuring operational excellence through modern deployment workflows
• Participating in on-call rotations for critical training jobs
You should have 4+ years of experience in data infrastructure, MLOps, or infrastructure engineering. Strong proficiency in Python is required, along with deep experience in Kubernetes-native tooling and debugging large-scale distributed systems. You're comfortable with modern columnar storage standards and solving data lake scalability challenges. This is a high-impact role in a rapidly growing AI company where you'll work on infrastructure problems at unprecedented scale.