SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mistral is a full-stack AI company providing frontier models, developer tools, applications, and compute infrastructure. They partner with enterprises across finance, manufacturing, defense, healthcare, and public sector to deploy customized AI systems.
You will join the compute infrastructure team as a Datacenter Hardware Engineer responsible for maintaining, troubleshooting, and scaling one of France's largest GPU/CPU clusters located in the Paris-area datacenter at Bruyères-le-Châtel. This is a hands-on, field-based role with direct impact on Mistral's ability to deliver breakthrough AI solutions.
Key responsibilities include:
• Diagnose and operate core server and cluster components, investigating and resolving compute/storage hardware issues (CPU, memory, drives, NICs, GPUs, PSUs) and interconnect problems (switches, cables, transceivers; Ethernet/InfiniBand). Perform safe interventions including power-off/lockout and ESD procedures to replace, re-seat, or recable components and restore service.
• Apply lockout/tagout (LOTO) and ESD discipline rigorously; follow pre/post-work checklists; maintain safe, tidy work areas.
• Perform first-line diagnostics using LEDs, POST, beep codes, and basic tests; capture evidence through photos and serial numbers; open, update, and close tickets with clear documentation.
• Contribute to preventive maintenance by providing feedback on proactive activities and monitoring; help convert ad-hoc checks into standard operating procedures, alerts, and dashboards.
• Manage parts and logistics: receive and track hardware, maintain labeled inventory, manage RMAs, and coordinate with vendors.
• Collaborate with senior hardware and firmware engineers on complex multi-node issues; communicate status and next steps clearly.
• Maintain current SOPs and checklists; ensure zero undocumented changes and audit-ready records.
You bring hands-on datacenter and server hardware experience: comfortable installing, re-seating, and swapping GPU/PCIe cards, NICs, PSUs, and drives in racks with proper cabling and labeling. You are disciplined and meticulous, following checklists and safety protocols without exception. You have practical electrical basics including power-off procedures, PPE awareness, and short-circuit risk awareness. You are comfortable working in racks with cooling, network, storage, and PDU systems; can safely lift and mount equipment within HSE limits. You communicate clearly with factual updates, are a reliable teammate, punctual, and process-minded.
Nice-to-have qualifications include HPC/AI/cloud production environment experience, large-fleet server installation and maintenance, basic networking (Ethernet/InfiniBand), basic Linux (boot/check), coding/automation skills (Python/Bash) for tooling and monitoring, inventory/RMA tool experience, and exposure to HPC/research/industrial environments.