SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer, AI Platform

Algolia - Paris, France - Hybrid - posted 2026-09-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: EUR 69,768 - 96,900 / annual

Algolia is the retrieval intelligence layer powering over 1.7 trillion queries annually for 18,000+ customers. The AI Platform team builds and operates shared production foundations supporting Algolia's evolving AI ecosystem, working at the intersection of Site Reliability Engineering, cloud infrastructure, software engineering, and AI. In this role, you will build and operate production infrastructure supporting AI-related workloads and services. You'll operate and improve highly available Kubernetes-based platforms, improve reliability through SLOs, observability, alerting and capacity management, and investigate production issues to turn findings into durable fixes. Your responsibilities span networking, databases, compute and service infrastructure. You'll improve CI/CD pipelines, deployment automation and developer experience, build and maintain infrastructure using Infrastructure as Code, participate in on-call and incident response, and collaborate with experienced engineers across the team while progressively taking ownership of broader production areas. The team emphasizes solving operational problems, automating repetitive work, and progressively taking ownership of complex systems at scale. You'll work in a high-trust environment with flexibility in how and where you work, though this position is based in Paris with hybrid-remote options depending on the role. REQUIREMENTS: - Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations - Strong experience with Infrastructure as Code and the lifecycle of cloud infrastructure - Solid experience building and operating CI/CD pipelines and automated deployment workflows - Hands-on experience with at least one major cloud provider: GCP, AWS or Azure - Good understanding of networking, distributed systems and reliability engineering - Experience with monitoring, observability and troubleshooting production systems - Strong automation mindset and ability to take ownership of well-defined production systems and progressively tackle more complex problems - Excellent written and spoken English NICE TO HAVE: - Go and/or Python engineering experience - Exposure to AI/ML infrastructure and inferences - Comfortable working AI-first, using coding agents, agentic workflows and AI-assisted debugging

Similar roles