SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mirantis is seeking a Senior Site Reliability Engineer to join its cloud development and operations team. You will contribute to designing, developing, and operating sophisticated cloud-based AI solutions built on the CNCF ecosystem, including Kubernetes, running on cutting-edge NVIDIA-certified hardware. This role focuses on deploying AI infrastructure, ensuring reliability, security, and performance of container infrastructure while mentoring team members and customers.
Key responsibilities include:
- Working with geographically distributed international teams on technical challenges and process improvements
- Developing, implementing, maintaining, and troubleshooting cloud and AI infrastructure solutions based on open source software
- Collaborating with stakeholders to gather and refine technical requirements
- Optimizing system performance, reliability, and scalability
- Troubleshooting, debugging, and resolving complex technical issues
- Participating in code reviews to maintain high quality standards
- Staying current with industry trends and best practices in cloud operations
- Designing and implementing AI-driven automation across the DevOps lifecycle
- Facilitating knowledge transfer to customers during delivery phases
You will work with an established leader in cloud infrastructure, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies. The role offers professional development, conference attendance, and a competitive compensation package.
Requirements:
- 5+ years of professional experience in DevOps with strong focus on cloud and infrastructure technologies, including Kubernetes and/or OpenStack
- Experience with high-performance data center processing, networking, and storage
- Exposure to Golang and working knowledge of other programming languages (Python, JavaScript)
- Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines
- Exceptional problem-solving and debugging skills across networking, storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security
- Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams
- Comfortable making independent judgment calls when working directly with customers with limited day-to-day oversight
- Excellent written and spoken English and customer-facing communication skills
- Commitment to innovation, continuous learning, and delivering high-quality results
- Ability to travel up to 25% if needed, including internationally
- Bachelor's degree in Computer Science or related field, or equivalent experience
Nice to have:
- Extensive experience in network and/or storage architecture
- Experience with high-performance computing or GPU infrastructure (GPU scheduling, MIG/vGPU, RDMA/RoCE, InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, NVIDIA AI Enterprise)
- Presence in open source community including upstream contributions and conference presentations
- Prior experience with commercial container and virtual compute infrastructure platforms (Rancher, OpenShift, VMware)