SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mind Robotics is building generalized physical AI systems for industrial deployment, starting with factory automation. The company develops robotic systems capable of dexterous, adaptive, and reasoning-intensive work in real-world environments.
As a DevOps Engineer, you'll build and operate the cloud infrastructure that enables the robotics, ML, and software teams to develop, test, deploy, and monitor robotic systems reliably at scale. You'll work across multiple teams to automate infrastructure, streamline deployments, improve developer productivity, and ensure robots can be continuously updated and monitored both in lab and production settings.
Key responsibilities include designing and maintaining scalable cloud infrastructure on AWS, GCP, or Azure; building Infrastructure as Code using Terraform; developing CI/CD pipelines for robotics software, embedded applications, and ML workflows; automating deployments across development robots, test environments, and production fleets; managing Kubernetes clusters and containerized services; implementing monitoring and observability using modern tools; improving system reliability, uptime, security, and disaster recovery; building internal developer tooling; partnering with robotics and firmware teams to create reproducible environments; supporting simulation infrastructure and large-scale testing; optimizing cloud costs; implementing security best practices; and participating in incident response and operational improvements.
Required qualifications: 4+ years in DevOps, Platform Engineering, SRE, or Infrastructure Engineering; strong Linux administration skills; cloud platform experience (AWS, GCP, or Azure); production Kubernetes and Docker experience; Infrastructure as Code expertise (Terraform preferred); CI/CD pipeline building experience (GitHub Actions, GitLab CI, Jenkins, Buildkite, or similar); configuration management and automation tools; Python, Go, or Bash proficiency; monitoring and logging implementation experience (Prometheus, Grafana, Datadog, ELK, OpenTelemetry); strong networking, security, authentication, and distributed systems knowledge; and production system troubleshooting experience.
Nice-to-have skills include robotics or autonomous systems support experience, fleet device deployment experience, GPU infrastructure and ML training cluster support, NVIDIA GPU/CUDA/distributed training knowledge, artifact repository management, simulation environment support, software supply chain security and SBOM knowledge, and edge computing or IoT deployment experience.