SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer (AI Platform)

ManyChat - Barcelona, Catalonia, Spain - Hybrid - posted 2026-07-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ManyChat is a leading Chat Marketing platform used by 1.5+ million customers worldwide, helping creators automate conversations across Instagram, Messenger, WhatsApp, and TikTok. The company has 350+ employees across three continents and is scaling rapidly with AI-powered features. You'll join as a Staff Site Reliability Engineer, taking ownership of infrastructure strategy and platform reliability. This is a high-impact individual contributor role where you'll bridge Infrastructure and Engineering teams, shaping how the platform scales and evolves. Key responsibilities include: - Maintain and harden AWS infrastructure (EC2, ALB/NLB, WAF, IAM, CloudWatch) - Operate and evolve EKS clusters powering Python-based AI services - Migrate existing services to Kubernetes using Terraform and Helm - Codify infrastructure with Terraform and manage host-level automation via Ansible - Build and improve CI/CD pipelines with GitHub Actions - Own observability: Prometheus, Grafana, alerting, and on-call readiness - Support OS-level patching, certificates, WAF rules, and infrastructure hygiene - Partner with engineers to guide best practices and drive platform reliability - Create clean, maintainable infrastructure documentation and playbooks - Support rare off-hours incidents Required qualifications: - 5+ years managing Linux in production (Ubuntu, Amazon Linux) - Strong Kubernetes experience (ideally EKS), Helm, and Terraform - Comfort running and debugging Python workloads in containers - Solid understanding of networking, IAM, and cloud security - Hands-on Nginx experience (Ingress and reverse proxy setups) - Excellent communication skills Nice-to-have: Advanced Ansible, PostgreSQL/RDS tuning, observability tools (Prometheus, Grafana, Loki), PHP production experience, TDD/CI-CD best practices, prior SRE exposure. The company offers hybrid onboarding with relocation support, comprehensive health insurance, professional development budget, flexible benefits, hybrid work, generous leave, in-office perks, and company-funded activities.

Similar roles