SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior DevOps Engineer

Dynamo AI - India - In-office - posted 2026-09-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Dynamo AI helps enterprises deploy AI systems that are reliable, secure, and production-ready. The company provides market-leading technical controls spanning AI evaluations, guardrails, agentic risk management, and observability, partnering with highly-regulated global organizations across financial services and government. We are seeking a Senior DevOps Engineer to build, scale, and operate the infrastructure powering our AI platform. This is a hands-on, high-ownership role requiring someone who can work independently, solve complex infrastructure problems, and thrive in a fast-paced startup environment. Key Responsibilities: - Design, build, and operate highly available production infrastructure on AWS, with strong expertise in EKS, EC2, VPC, S3, RDS/Aurora, IAM, ECR, ElastiCache, and Load Balancers - Build and improve CI/CD and release automation using Jenkins, GitHub Actions, Helm, ArgoCD, and GitOps - Manage infrastructure using Terraform and Infrastructure as Code principles - Build and operate Kubernetes platforms at production scale, including cluster management, upgrades, autoscaling, networking, security, and troubleshooting - Run and operate AI/ML workloads in production with deep understanding of infrastructure challenges for AI systems - Deploy, scale, monitor, and optimize AI inference and model-serving workloads across Kubernetes and cloud infrastructure - Work with GPU-based workloads, including GPU scheduling, utilization, autoscaling, capacity planning, and optimization - Drive infrastructure efficiency by balancing performance, reliability, scalability, and cost across AI workloads - Own monitoring, logging, and observability using Prometheus, Grafana, Thanos, and OpenTelemetry - Implement secure secrets management using HashiCorp Vault and External Secrets Operator - Develop automation and internal tooling using Python and Bash - Drive improvements around reliability, security, scalability, performance, and infrastructure cost - Participate in production incidents, root-cause analysis, and drive long-term fixes - Work closely with Engineering, ML/AI, Security, and Product teams You should be someone who builds, automates, and takes ownership—not someone who simply operates existing infrastructure. You must be comfortable with ambiguity, willing to dive deep into production problems, and constantly looking for ways to make the platform more reliable, secure, scalable, and cost-efficient. Understanding that AI infrastructure has different operational challenges is critical; you should help run AI systems efficiently at scale, making the right trade-offs between GPU utilization, performance, reliability, scalability, and cost. Requirements: - 5+ years of strong hands-on experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure - Strong production experience with AWS and Kubernetes/EKS - Proven experience running AI/ML workloads or GPU-based workloads in production (highly valuable) - Strong understanding of CI/CD, Infrastructure as Code, GitOps, observability, and cloud security - Excellent scripting and automation skills in Python and Bash - Experience operating production systems and troubleshooting complex infrastructure issues independently - Strong understanding of scaling, performance optimization, resource utilization, and cost management, particularly for compute-intensive workloads - Strong ownership mindset with the ability to take a problem from design to production - Experience working in a startup or fast-moving engineering environment (highly valued) - Strong communication skills and ability to work effectively across teams Nice to Have: - Experience with AI inference platforms, model serving, LLM infrastructure, or ML platforms - Experience with GPU infrastructure such as NVIDIA GPUs and Kubernetes GPU scheduling - Experience with multi-region or highly distributed systems - Experience with SOC 2, ISO 27001, or other security/compliance requirements - Experience with PostgreSQL, MongoDB, Redis, Kafka, or similar distributed systems

Similar roles