SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mistral is building full-stack AI solutions—from frontier models to developer tools, applications, and compute infrastructure. The company partners with enterprises across finance, manufacturing, defense, healthcare, and the public sector to co-create customized AI systems.
As a Site Reliability Engineer on the Cloud Platform team, you will own the reliability, scalability, and performance of Mistral's cloud infrastructure and customer-facing applications. You'll work closely with software engineers and product teams to ensure systems meet expectations for both internal and external customers.
Key responsibilities include:
- Design, build, and maintain scalable, highly available, fault-tolerant infrastructure supporting the cloud platform
- Operate production systems, troubleshoot issues, respond to incidents, and manage infrastructure scaling
- Implement and improve monitoring, alerting, and incident response systems to minimize downtime
- Develop and maintain CI/CD workflows, containerization, orchestration, monitoring, and logging tools
- Participate in on-call rotations, perform root cause analysis, and drive continuous improvement in automation
- Collaborate with software engineers to enable safe, reproducible model-training experiments
- Build cloud platform abstractions that simplify infrastructure for science and engineering teams
- Ensure infrastructure adheres to security best practices and compliance requirements
- Document processes and procedures for consistency and knowledge sharing
You bring 5+ years of DevOps or SRE experience with strong expertise in bare metal infrastructure and distributed systems. You have hands-on experience with site reliability issues, production troubleshooting, and on-call rotations. You're proficient with CI/CD, Docker, Kubernetes, infrastructure-as-code (Terraform/CloudFormation), and monitoring tools (Prometheus, Grafana, ELK, Datadog). You code fluently in Python, Go, or Bash, understand networking and security concepts, and have excellent problem-solving and communication skills. Experience with AI/ML environments, HPC systems, or modern AI-oriented solutions (Fluidstack, Coreweave, Vast) is a plus. A Master's degree in Computer Science, Engineering, or related field is required.