SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Ivo is an AI company that builds products to help enterprises solve complex contracting problems using large language models. The company has achieved several infrastructure and product firsts, including AI agents embedded in MS Word, agentic RAG systems, large-scale LLM-based legal fact extraction, and advanced contract analysis capabilities.
As a Staff Infrastructure Engineer, you will build and own the foundational infrastructure that powers Ivo's products. This is not a "keep the lights on" role—you'll design and operate the systems that enable the entire company to function at scale.
Key Responsibilities:
- Run and own multi-cluster, multi-region Kubernetes deployments across AWS, GCP, and Azure with robust failover and disaster recovery
- Design infrastructure strategy for isolated workloads, balancing the different performance and cost needs of ML versus API traffic
- Build internal tools and platforms that enable consistent, reliable deployments from development to production
- Own CI/CD pipelines using GitHub Actions and infrastructure-as-code with Pulumi
- Implement security controls (RBAC, workload identity, secrets management, data residency, audit trails) that are transparent to engineers but critical to customers
- Define and maintain honest SLOs/SLIs that prevent alert fatigue while ensuring reliability
- Improve observability across the entire engineering environment to detect bugs and performance issues early
- Automate repetitive tasks and optimize build times and developer workflows
- Lead incident response and postmortem processes that drive organizational learning
- Work deeply embedded with the engineering team and contribute to pushing the frontiers of LLM infrastructure
You'll be working with a team of high-caliber engineers who move fast, maintain high standards, and care deeply about their craft. The role offers significant responsibility and impact, with the opportunity to shape how a fast-growing AI company operates at scale.
Requirements:
- 8+ years of experience in DevOps/infrastructure engineering
- Deep expertise in Kubernetes (not just deployment-level knowledge)
- Hands-on experience with infrastructure-as-code (Pulumi or Terraform) and CI/CD (GitHub Actions, Docker)
- Proven track record managing multi-cluster, multi-region deployments
- Strong debugging skills and systematic approach to root-cause analysis
- Understanding of how compliance and legal requirements translate into engineering constraints
- Full-stack range sufficient to unblock yourself and not get stuck
- Ability to think in failure modes and design for resilience from day one
- Experience in a startup environment preferred but not required
- Relentless resourcefulness, strong sense of urgency, and bias toward action
- Excitement about building AI products and the adventure of scaling a company