SlipstreamJobsFresh Startup & VC-Backed Jobs

AI Infrastructure Systems Engineer (Amsterdam & London)

Together AI - Amsterdam, North Holland, Netherlands - Hybrid - posted 2025-04-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Together AI is building one of the world's largest GPU fleets for frontier model training and inference. As an AI Infrastructure Systems Engineer, you'll design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. You'll develop AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. You'll create Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Your work will focus on maximizing GPU availability, utilization, performance, and reliability across thousands of accelerators. Key responsibilities include building automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. You'll create internal platforms and developer tools that allow infrastructure to be managed through software rather than manual operations. You'll continuously improve deployment velocity, reliability, and operational efficiency through automation, partnering closely with hardware, networking, platform, and AI teams. Required qualifications: 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking and passion for solving complex infrastructure challenges through software. An automation-first mindset—if a task is repeated, your instinct is to build a system to eliminate it. Bonus experience includes GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, InfiniBand or RoCE networking, bare-metal provisioning, large-scale AI training/inference clusters, hardware health monitoring, distributed storage systems, and AI agents for autonomous infrastructure operations. Together AI is a research-driven AI company focused on lowering the cost of modern AI systems through co-designed software, hardware, algorithms, and models. The team has contributed to leading open-source research including FlashAttention, Hyena, FlexGen, and RedPajama.

About Together AI

AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.

Similar roles