SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

Mistral - Paris, Île-de-France, France - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Mistral is building full-stack AI solutions, from frontier models to developer tools and compute infrastructure. As a Site Reliability Engineer on the Platform team, you will own the reliability, scalability, and performance of Mistral's platform and customer-facing applications. You'll balance day-to-day production operations with long-term infrastructure improvements. Your responsibilities include designing and maintaining highly available, fault-tolerant systems for web services, inference environments, and ML workloads. You'll implement monitoring, alerting, and incident response systems; develop CI/CD and infrastructure-as-code workflows; and participate in on-call rotations for production support. Key focus areas: ensuring seamless replication across HPC clusters, operating production systems, troubleshooting infrastructure issues, driving automation improvements with Kubernetes and Terraform, and collaborating with AI/ML research teams to enable reproducible experiments. You'll also work with security teams on compliance and build cloud-agnostic abstractions for science and engineering teams. Required: Master's degree in Computer Science or related field; 7+ years in DevOps/SRE roles with strong cloud and distributed systems expertise. You need hands-on experience with site reliability issues, root cause analysis, and on-call operations. Technical proficiency required in CI/CD, Docker, Kubernetes, infrastructure-as-code (Terraform/CloudFormation), monitoring tools (Prometheus, Grafana, ELK, Datadog), and scripting (Python, Go, Bash). Experience with HPC systems or AI-oriented compute platforms (Fluidstack, Coreweave, Vast) is a plus.

Similar roles