SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mirantis is seeking a Senior Site Reliability Engineer to join its cloud development and operations team. You will contribute to designing, developing, and operating sophisticated cloud-based AI solutions built on the CNCF ecosystem, including Kubernetes, running on cutting-edge NVIDIA-certified hardware. This role focuses primarily on deploying AI infrastructure, following architecture and implementation designs produced by the engineering team.
You will play a pivotal role in ensuring the reliability, security, and performance of container infrastructure while mentoring team members and Mirantis customers. Working closely with geographically distributed international teams, you will define technical strategies, solve complex challenges, and ensure seamless integration of cloud and software services.
Key Responsibilities:
- Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software
- Optimize system performance, reliability, and scalability across distributed systems
- Troubleshoot, debug, and resolve complex technical issues in networking, storage, and Kubernetes environments
- Design and implement AI-driven automation across the DevOps lifecycle
- Collaborate with stakeholders to gather and refine technical requirements
- Participate in code reviews to maintain high quality standards
- Facilitate knowledge transfer to customers during delivery phases
- Stay current with industry trends and best practices in cloud operations
- Work with geographically distributed teams on technical challenges and process improvements
- Travel up to 25% as needed, including internationally
Requirements:
- 5+ years of professional experience in DevOps with strong focus on cloud and infrastructure technologies, including Kubernetes and/or OpenStack
- Experience with high-performance data center processing, networking, and storage
- Exposure to Golang and working knowledge of other programming languages (Python, JavaScript)
- Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines
- Exceptional problem-solving and debugging skills across networking, storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security
- Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams
- Comfortable making independent judgment calls when working directly with customers with limited day-to-day oversight
- Excellent written and spoken English and customer-facing communication skills
- Bachelor's degree in Computer Science or related field, or equivalent experience
- Commitment to innovation, continuous learning, and delivering high-quality results
Nice to Have:
- Extensive experience in network and/or storage architecture
- Experience with high-performance computing or GPU infrastructure (GPU scheduling, MIG/vGPU, RDMA/RoCE, InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, NVIDIA AI Enterprise)
- Presence in open source community including upstream contributions and conference presentations
- Prior experience with commercial container and virtual compute infrastructure platforms (Rancher, OpenShift, VMware)