SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Algolia is a market leader in AI Search, powering over 30 billion search requests weekly for 18,000+ businesses including Under Armour, Stripe, and Walgreens. The company raised $150M in Series D funding at a $2.25B valuation.
The Infrastructure as a Service (IaaS) team is leading one of Algolia's most significant engineering transformations. For years, Algolia has operated approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability customers expect. The team is now building a unified cloud and Kubernetes platform to support future growth—not a simple lift-and-shift, but a fundamental rethinking of how Algolia provisions, secures, operates, observes, upgrades, and scales production infrastructure.
As a Site Reliability Engineer (P3) in IaaS, you will be a hands-on contributor helping build the next generation of Algolia's production infrastructure. You will develop deep expertise in cloud, Kubernetes, reliability, and automation while contributing across cloud foundations, lifecycle capabilities, and large-scale production operations.
Key responsibilities include:
- Building and improving Cloud Baseline capabilities: identity and access, networking, security, resource inventory, tagging, and auditability
- Developing and maintaining infrastructure as code and automation for cloud environments and Kubernetes
- Contributing to reliable, repeatable cloud and cluster lifecycle operations
- Building self-service capabilities, reusable modules, and documentation that make the safe path easy for platform consumers
- Reducing manual work and configuration drift through automation, testing, GitOps practices, and standardization
- Leveraging automation and AI-assisted engineering tools to improve infrastructure analysis, documentation, and safe changes
- Improving observability, monitoring, alerting, capacity management, and operational documentation
- Investigating production issues, participating in on-call rotation, and turning lessons learned into lasting improvements
- Collaborating with Infrastructure, Security, FinOps, and engineering teams to deliver reliable, secure, and cost-aware platform capabilities
Algolia values autonomy, diversity, and collaboration. The company operates with a flexible workplace model emphasizing individual impact and contribution over physical location, with offices in Paris, NYC, London, Sydney, and Bucharest, plus remote options.
REQUIREMENTS:
- Hands-on production knowledge of AWS or GCP
- Practical Kubernetes knowledge and interest in operating it in production
- Familiarity with infrastructure as code, ideally Terraform
- Programming or scripting skills in Python, Go, or equivalent language
- Strong Linux and networking fundamentals
- Strong interest in reliability, automation, and solving production problems
- Comfort adopting AI-assisted engineering tools with sound judgment for critical production systems
- Ability to communicate clearly and work effectively with distributed teams
- Excellent spoken and written English skills
NICE TO HAVE:
- Familiarity with multiple public cloud providers
- Knowledge of GitOps or policy-as-code tooling (Argo CD, Helm, OPA, Kyverno)
- Experience with cloud migration, platform engineering, or large-scale infrastructure transformation