SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lio is building an AI workforce platform for enterprise procurement. As a Site Reliability Engineer, you will work directly with the CTO to design, build, and operate the infrastructure powering one of Europe's fastest-growing enterprise AI startups.
You will own critical infrastructure responsibilities including designing and operating highly reliable cloud infrastructure for production AI workloads, scaling the platform globally with multi-region deployments and high-availability architectures, and optimizing databases for agentic AI workloads including Vector Search, Hybrid RAG, and MCP. You'll improve backend performance across latency, throughput, and resource efficiency, build observability, SLOs, alerting, and incident response processes, and automate deployments, infrastructure, and developer workflows.
Your technical foundation should include production experience operating workloads on major cloud platforms (AWS, GCP, or Azure), strong Python skills with backend service optimization experience, solid understanding of distributed systems and asynchronous processing, and hands-on experience with monitoring and observability tools. Database optimization at scale is important—MongoDB experience is a plus. You should be comfortable with CI/CD pipelines, preferably GitHub Actions, and have a passion for automation and solving complex scaling challenges.
This is a 100% on-site role in the Munich office where you'll partner closely with product engineering teams and help shape infrastructure decisions that define the company's trajectory. You'll work on challenging problems around global scaling, reliability, and agentic AI infrastructure.