SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Simple AI builds state-of-the-art voice AI agents for enterprise customers, handling real phone operations like customer support, order intake, and lead qualification for companies including DoorDash, xAI, and Omaha Steaks. The company raised $14M from top-tier investors including Y Combinator and has grown to 20+ engineers.
You'll join as a Platform Engineer working directly with the founders to build and operate the infrastructure behind real-time voice conversations. This is a hands-on role with meaningful ownership over infrastructure architecture and deployment practices.
Key responsibilities:
- Build and operate AWS infrastructure using Terraform and Kubernetes, including EKS clusters, networking, IAM, load balancing, and secrets management
- Develop reusable infrastructure modules and maintain consistent development, staging, and production environments
- Build deployment workflows with automated checks, clear visibility, and reliable rollbacks
- Design autoscaling and graceful shutdown behavior for services handling long-running voice conversations
- Improve observability across infrastructure and application services with dashboards, alerts, and reliability metrics
- Investigate production issues, improve incident response, and address recurring problems at their source
- Partner with application and AI engineers to deploy new services, improve performance, and manage costs
- Make it easier for engineers to provision resources, deploy changes, and debug services safely
The role demands deep expertise in real-time systems where latency is immediately noticeable, calls are stateful, and deployments must happen without interrupting customers.
REQUIREMENTS:
- Hands-on production experience running Kubernetes and AWS infrastructure
- Strong Terraform experience, including maintainable modules, state management, and safe infrastructure changes
- Solid understanding of Linux, containers, networking, and distributed systems
- Experience building CI/CD pipelines and operating services through deployments, incidents, and growth
- Ability to write software and automation in Python, Go, or TypeScript
- Good judgment about infrastructure decisions and when to keep things simple
- Ownership mindset, clear communication, and follow-through on lasting fixes
NICE TO HAVE:
- Experience with Argo CD, Helm, or GitOps tooling
- Familiarity with real-time audio, SIP, WebRTC, or latency-sensitive systems
- Experience operating PostgreSQL, Redis, queues, and background workers
- Experience deploying GPU workloads or self-hosted AI inference
- Infrastructure security and SOC 2 compliance experience