SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Maven AGI is an enterprise AI platform founded in July 2023 by executives from HubSpot, Google, and Stripe. The company builds conversational AI agents for autonomous customer support at scale, unifying fragmented systems and enabling intelligent actions without costly infrastructure changes. The team includes talent from Google, Meta, Amazon, Microsoft, and Stripe.
As Senior DevOps Engineer, you will own and evolve the infrastructure powering Maven AGI's AI platform. You'll design, build, and operate production systems across cloud providers (Azure, AWS), Kubernetes clusters, on-premises environments, and CI/CD pipelines. Your work directly impacts platform availability, developer velocity, and customer trust as the company onboards enterprise customers with complex deployment and security requirements.
Key responsibilities include: designing and maintaining cloud and on-premise infrastructure using infrastructure-as-code tools (Pulumi, Bicep, Terraform); owning Kubernetes cluster operations including deployments, scaling, monitoring, and incident response; building and optimizing CI/CD pipelines for large-scale monorepos; implementing observability across services (metrics, logging, tracing, alerting); driving reliability practices including SLOs, capacity planning, disaster recovery, and runbook development; operationalizing enterprise AI deployments on-premise with GPU resource orchestration and model inference performance tuning; collaborating with engineering teams to improve developer experience; managing secrets, access controls, and infrastructure security; and evaluating new tooling to reduce operational toil.
Required qualifications: 3-7 years of professional DevOps/SRE/Infrastructure experience; deep expertise with Kubernetes in production (AKS, EKS, or GKE); strong infrastructure-as-code skills; experience operating CI/CD systems; proficiency in scripting/programming languages (Python, Go, TypeScript, Bash); solid understanding of IaaS providers, networking, DNS, load balancing, and TLS; experience with monitoring and observability stacks; experience with multi-cloud or hybrid deployments; strong communication and cross-team collaboration skills; and comfort operating in fast-paced startup environments.
Nice-to-have skills include experience with GPU infrastructure and ML/LLM serving workloads (vLLM, TEI), familiarity with Temporal or workflow orchestration systems, security and compliance background (SOC 2, HIPAA, GDPR), and cost optimization experience at scale.