SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mirantis is seeking a Senior Site Reliability Engineer to join its cloud development and operations team. You will contribute to designing, developing, and operating sophisticated cloud-based AI solutions built on the CNCF ecosystem, including Kubernetes, running on cutting-edge NVIDIA-certified hardware. This role focuses primarily on deploying AI infrastructure, ensuring reliability, security, and performance of container infrastructure while mentoring team members and customers.
Key responsibilities include:
- Working with geographically distributed international teams on technical challenges and process improvements
- Developing, implementing, maintaining, and troubleshooting cloud and AI infrastructure solutions based on open source software
- Collaborating with stakeholders to gather and refine technical requirements
- Optimizing system performance, reliability, and scalability
- Troubleshooting, debugging, and resolving complex technical issues
- Participating in code reviews to maintain high quality standards
- Staying current with industry trends and best practices in cloud operations
- Designing and implementing AI-driven automation across the DevOps lifecycle
- Facilitating knowledge transfer to customers during delivery phases
You will work closely with stakeholders to define technical strategies, solve complex challenges, and ensure seamless integration of cloud and software services. This role offers the opportunity to make significant impact while driving innovation in a rapidly evolving cloud ecosystem.
Requirements:
- 5+ years of professional experience in DevOps with strong focus on cloud and infrastructure technologies, including Kubernetes and/or OpenStack
- Experience with high-performance data center processing, networking, and storage
- Exposure to Golang and working knowledge of other programming languages (Python, JavaScript)
- Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines
- Exceptional problem-solving and debugging skills across networking, storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security
- Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams
- Comfortable making independent judgment calls when working directly with customers with limited day-to-day oversight
- Excellent written and spoken English
- Excellent customer-facing communication skills
- Commitment to innovation, continuous learning, and delivering high-quality results
- Ability to travel up to 25% if needed, including internationally
- Bachelor's degree in Computer Science or related field, or equivalent experience
Nice to have:
- Extensive experience in network and/or storage architecture
- Experience with high-performance computing or GPU infrastructure (GPU scheduling, MIG/vGPU, RDMA/RoCE, InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, NVIDIA AI Enterprise)
- Presence in open source community including upstream contributions and conference presentations
- Prior experience with commercial container and virtual compute infrastructure platforms (Rancher, OpenShift, VMware)