SlipstreamJobsFresh Startup & VC-Backed Jobs

Technical Support Engineer (L2) - Compute

Mistral - Paris, Île-de-France, France - Hybrid

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Mistral is building a new Compute Support team to ensure the reliability, performance, and scalability of GPU clusters in Kubernetes environments. As a founding member of this team, you will shape its processes, standards, and culture while providing technical support for AI workloads at scale. You will serve as both the first point of contact for customers and internal teams on compute-related issues, and as the L2 escalation point for complex technical problems. This hybrid L1/L2 support role blends system administration, customer-facing communication, and compute infrastructure expertise. Key responsibilities include: **First Point of Contact**: Act as the primary interlocutor for customers and internal teams, triaging and prioritizing incoming requests while meeting SLAs for acknowledgment, first response, and resolution. Gather and analyze issue details (logs, error messages, system metrics) to diagnose problems efficiently and provide clear, actionable guidance including temporary workarounds. **Technical Support**: Serve as L2 escalation for Linux, Kubernetes, and compute infrastructure issues, providing deep troubleshooting for GPU/TPU workloads (CUDA errors, memory leaks, job failures), Kubernetes clusters (pod crashes, node failures, networking misconfigurations), and bare metal/cloud environments (AWS EC2, GCP VMs, HPC clusters). Diagnose performance bottlenecks, hardware failures, and resource contention in distributed systems. Analyze system metrics, logs, and traces using tools like dmesg, journalctl, nvidia-smi, Prometheus, and Grafana. Participate in on-call rotations for 24/7 critical system support. Optimize system configurations for HPC and AI workloads. **Kubernetes & Containerization**: Debug Kubernetes clusters focusing on pod/node issues (CrashLoopBackOff, OOMKilled, ImagePullBackOff), networking and storage (CNI plugins, PersistentVolumes, StorageClasses), and resource management (Requests/Limits, QOS classes, node affinity). Troubleshoot container runtime issues (Docker, containerd). Collaborate with SRE teams on IaC (Terraform, Ansible) and Go-based tooling. **Documentation & Process Improvement**: Create and maintain runbooks, playbooks, and internal documentation for common compute issues. Contribute to post-mortems with actionable follow-ups. Train internal teams on compute infrastructure best practices. Help define and refine support processes as a founding team member. **Requirements**: - 5+ years of experience in system administration, technical support, or infrastructure operations, with strong focus on compute-heavy environments (bare metal, cloud, HPC, or virtualization) - Deep Linux/Unix expertise: kernel-level debugging (OOM killer, I/O bottlenecks, CPU throttling), performance tuning (sysctl, ulimit, filesystem optimizations), networking troubleshooting (iptables, tcpdump, DNS, NFS, SSH), storage management (LVM, RAID, NVMe, GPU-local storage) - Mandatory hands-on Kubernetes experience: debugging pods, nodes, and clusters (kubectl describe, kubectl logs, crictl), networking and storage (CNI plugins, PersistentVolumes), resource management (Requests/Limits, QOS classes) - Familiarity with containerization (Docker, containerd) and basic IaC awareness - Proficiency with monitoring tools (Prometheus, Grafana, ELK, OpenTelemetry) - Basic scripting/automation skills (preferably Go, but Bash or Python acceptable for ad-hoc tasks) - Strong plus: experience with bare metal, virtualization, or HPC environments (Fluidstack, Coreweave, Vast) - Excellent problem-solving and communication skills; ability to explain technical issues clearly to technical and non-technical stakeholders - Customer-focused mindset with passion for efficient issue resolution and system reliability - Ability to work on-call rotations and handle high-pressure situations with structured, calm approach - Commitment to meeting tight SLAs for issue acknowledgment, triage, and first response

Similar roles