SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Together AI is building an AI-native cloud platform combining fast LLM inference with state-of-the-art infrastructure. The Together Cloud Infrastructure team develops the GPU Clusters product—a self-serve cloud console delivering high-performance, AI-ready GPU clusters as the virtualized infrastructure layer powering inference, reinforcement learning, and fine-tuning services.
As a Senior Software Engineer, you will architect and build the next-generation AI cloud platform: a highly available, globally distributed infrastructure virtualizing cutting-edge ML hardware (GB200s/GB300s, BlueField DPUs). You'll enable ML practitioners with self-serve services including on-demand and managed Kubernetes/Slurm clusters, serving both internal SaaS products and external cloud customers across dozens of data centers worldwide.
Key responsibilities include:
- Design, build, and maintain performant, secure, highly-available backend services and operators automating hardware management (Infiniband partitioning, parallel storage provisioning, VM provisioning)
- Build the IaaS software layer for new GB200 data centers with thousands of GPUs
- Design distributed GPU scheduling and global management planes powering on-demand clusters across multiple data centers
- Develop infrastructure supporting internal inference, RL, and fine-tuning products plus external customers
- Build systems automating capacity scaling and new capacity onboarding
- Work on a global multi-exabyte high-performance object store for pretraining datasets and model weights
- Build advanced observability stacks with automated node lifecycle management for fault-tolerant distributed training and large-scale inference
- Perform architecture and research on decentralized AI workloads
- Contribute to the open-source Together AI platform
- Create services, tools, developer documentation, and testing frameworks for robustness
You bring 5+ years of professional software development with proficiency in backend languages (Golang preferred) and 5+ years building high-performance, production-quality code with demonstrated ownership of large-scale projects. You have hands-on experience with high-performance and globally distributed microservice architectures across cloud providers (AWS, Azure, GCP). Strong systems knowledge across compute, networking, and storage is essential, along with experience building and operating reliable, customer-facing production systems using infrastructure automation (Terraform, Ansible), observability stacks (Prometheus, Grafana), and CI/CD pipelines.
Preferred qualifications include deep Kubernetes internals experience (operators, plugins, custom schedulers), VM/hypervisor expertise (QEMU/KVM, cloud-hypervisor, VFIO, virtio), datacenter networking (VLAN, VXLAN, VPN, VPC, OVS/OVN), Cluster API, high-performance compute/networking/storage, GPU/Infiniband virtualization, IaaS/PaaS systems at scale, DPU/SmartNIC experience, and GPU programming knowledge (NCCL, CUDA).
About Together AI
AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.