SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Together Cloud Infrastructure

Together AI - San Francisco, CA, United States - In-office - posted 2025-06-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Together AI is building an AI-native cloud platform combining fast LLM inference with state-of-the-art infrastructure. The Together Cloud Infrastructure team develops the GPU Clusters product—a self-serve cloud console delivering high-performance, AI-ready GPU clusters as the virtualized infrastructure layer powering inference, reinforcement learning, and fine-tuning services. As a Senior Software Engineer, you will architect and build the next-generation AI cloud platform: a highly available, globally distributed infrastructure virtualizing cutting-edge ML hardware (GB200s/GB300s, BlueField DPUs). You'll enable ML practitioners with self-serve services including on-demand and managed Kubernetes/Slurm clusters, serving both internal SaaS products and external cloud customers across dozens of data centers worldwide. Key responsibilities include: - Design, build, and maintain performant, secure, highly-available backend services and operators automating hardware management (Infiniband partitioning, parallel storage provisioning, VM provisioning) - Build the IaaS software layer for new GB200 data centers with thousands of GPUs - Design distributed GPU scheduling and global management planes powering on-demand clusters across multiple data centers - Develop infrastructure supporting internal inference, RL, and fine-tuning products plus external customers - Build systems automating capacity scaling and new capacity onboarding - Work on a global multi-exabyte high-performance object store for pretraining datasets and model weights - Build advanced observability stacks with automated node lifecycle management for fault-tolerant distributed training and large-scale inference - Perform architecture and research on decentralized AI workloads - Contribute to the open-source Together AI platform - Create services, tools, developer documentation, and testing frameworks for robustness You bring 5+ years of professional software development with proficiency in backend languages (Golang preferred) and 5+ years building high-performance, production-quality code with demonstrated ownership of large-scale projects. You have hands-on experience with high-performance and globally distributed microservice architectures across cloud providers (AWS, Azure, GCP). Strong systems knowledge across compute, networking, and storage is essential, along with experience building and operating reliable, customer-facing production systems using infrastructure automation (Terraform, Ansible), observability stacks (Prometheus, Grafana), and CI/CD pipelines. Preferred qualifications include deep Kubernetes internals experience (operators, plugins, custom schedulers), VM/hypervisor expertise (QEMU/KVM, cloud-hypervisor, VFIO, virtio), datacenter networking (VLAN, VXLAN, VPN, VPC, OVS/OVN), Cluster API, high-performance compute/networking/storage, GPU/Infiniband virtualization, IaaS/PaaS systems at scale, DPU/SmartNIC experience, and GPU programming knowledge (NCCL, CUDA).

About Together AI

AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.

Similar roles