SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff/Principal DevOps Engineer, AI Inference

Lila - Cambridge, MA, United States - In-office - posted 2026-07-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lila is seeking a Staff/Principal DevOps Engineer to design and optimize infrastructure for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, focusing on building systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will own the design and implementation of GPU/accelerator infrastructure on Kubernetes, including scheduling, resource isolation, multi-tenant GPU sharing, and topology-aware placement for inference workloads. You'll build model serving platforms using frameworks such as vLLM, Triton Inference Server, and TGI, with optimized batching, caching, and request routing. Key responsibilities include implementing intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium), designing autoscaling systems that dynamically match inference compute supply with demand, and building production-grade deployment pipelines for ML models with canary rollouts, A/B testing, and safe rollback across multi-region deployments. You will manage infrastructure-as-code using Terraform and Helm for GPU-accelerated EKS clusters, implement comprehensive observability and performance optimization including GPU utilization monitoring and inference latency profiling, and design CI/CD pipelines for model artifacts with container image builds and automated inference benchmarking. Additional focus areas include AWS cloud infrastructure optimization for ML workloads, cost optimization through right-sizing and spot instance strategies, and capacity planning for rapidly scaling AI systems. Required expertise includes deep experience with DevOps, SRE, or Platform Engineering operating GPU/accelerator infrastructure at scale, Kubernetes for ML workloads with GPU scheduling and resource management, AWS infrastructure-as-code with hands-on GPU compute experience, model serving infrastructure knowledge, networking for distributed inference, and strong Python proficiency for automation and ML framework integration. Bonus qualifications include LLM inference optimization experience, hands-on work with multiple accelerator families, multi-region deployment expertise, Rust or Go proficiency, SRE practices for ML systems, model registry and artifact versioning experience, custom observability platform building, and prior startup/high-growth experience.

Similar roles