SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mistral is seeking an experienced Site Reliability Engineer to shape the reliability, scalability, and performance of its AI platform and customer-facing applications. You will work with software engineers and research teams to ensure systems meet internal and external customer expectations.
In the operations domain, you will design, build, and maintain scalable, highly available, and fault-tolerant infrastructures supporting web services and ML workloads. You'll ensure platform, inference, and model training environments remain highly available across multiple HPC clusters. You'll operate production systems, troubleshoot issues, implement monitoring and alerting systems, and maintain CI/CD, containerization, and orchestration workflows. You'll participate in on-call rotations and perform root cause analysis on incidents.
On the development side, you'll drive continuous improvement in infrastructure automation using Kubernetes, Flux, and Terraform. You'll collaborate with AI/ML researchers to develop reproducible model-training solutions and build a cloud-agnostic platform abstracting science from infrastructure. You'll design workflows and tooling to improve system reliability, work with the security team on compliance, document processes, and contribute to open-source projects and research.
Required qualifications include a Master's degree in Computer Science or related field, 7+ years of DevOps/SRE experience, strong cloud computing and distributed systems knowledge, exposure to critical production environments, experience with reliability KPIs, hands-on CI/CD and Kubernetes expertise, monitoring/logging tools proficiency (Prometheus, Grafana, ELK, Datadog), infrastructure-as-code tools (Terraform, CloudFormation), scripting languages (Python, Go, Bash), and strong networking and security fundamentals.
Preferred experience includes AI/ML environments, high-performance computing systems (Slurm), and modern AI-oriented cloud solutions (Fluidstack, Coreweave, Vast). The role offers comprehensive benefits including healthcare, parental leave, retirement plans, relocation support, and wellness programs.