SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mirantis is building an enterprise AI infrastructure product that enables organizations to run and govern large language models on their own Kubernetes clusters. You will join a small senior team early in the product lifecycle with broad ownership of the model-serving layer and its path to production.
You will design and build LLM serving infrastructure on Kubernetes, handling deployment, GPU scheduling, scaling, and model lifecycle management. You'll package the platform for enterprise environments using Helm-based installs, upgrades, and support for restricted/offline networks. Your work will integrate the serving layer with the platform's API gateway, identity, and metering services, and you'll build observability solutions for operating GPU inference in production, including serving metrics and GPU telemetry. You'll contribute across a multi-service codebase and help set engineering direction through design documents and code reviews.
This is a remote-first role on a small team that values written communication and high autonomy. You'll work with an established Silicon Valley leader in cloud infrastructure, collaborating with talented colleagues serving Fortune 500 and Global 2000 customers on next-generation cloud technologies.
REQUIREMENTS:
- 5+ years of software engineering experience in infrastructure, platform, or distributed systems
- Deep hands-on Kubernetes experience: building and operating production workloads and Helm charts (not just consuming managed clusters)
- Experience with GPU workloads or LLM inference, OR strong adjacent systems experience with a track record of learning fast
- Strong Go programming skills
- Solid CI/CD and infrastructure-as-code skills
- Fluency with AI-assisted development tools (Claude Code, OpenAI Codex) as part of daily engineering workflow
- Comfortable with high autonomy on a small, remote-first, written-culture team
NICE TO HAVE:
- Inference performance work (quantization, batching, caching) or distributed serving frameworks
- Enterprise deployment experience (air-gapped installs, SSO/OIDC, supply-chain security)
- UI development experience (React/TypeScript)
- Open-source contributions in Kubernetes or ML-infrastructure ecosystems