SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 260,000 - 300,000 / annual
Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle. The Together Cloud team operates the flagship GPU Clusters IaaS product, providing high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering inference, reinforcement learning, and fine-tuning products.
As a Staff Software Engineer focusing on AI Compute, you will set technical direction for and build major components of the next-generation AI cloud platform—a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware (GB300s/VRs, BlueField DPUs, InfiniBand, and dual/quad-plane RoCEv2 fabrics). This virtualized computing platform powers Together's own SaaS products and serves external cloud customers through self-serve offerings such as on-demand and reserved Kubernetes/Slurm clusters across dozens of data centers and hundreds of thousands of GPUs.
This is an architect-and-build role. You will own the GPU and network virtualization stack (hypervisor, kernel, and SDN work) that keeps GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware. You will architect and roadmap the in-DC IaaS layer—services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in data centers (VMs, parallel filesystems, VPCs, InfiniBand partitions). You will lead the build-out for a new Vera Rubin data center with thousands of GPUs, from hardware bring-up to customer-facing API.
You will design the GPU scheduling and global management plane—the distributed control plane behind on-demand and reserved clusters across dozens of data centers, including systems that scale per-cluster limits and automate onboarding of new capacity. You will architect monitoring and automated remediation for fault tolerance, defining the strategy for automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference running through hardware failures.
Beyond individual contribution, you will set technical direction across teams by leading design reviews, resolving cross-cutting architectural disagreements, unblocking cross-team dependencies and integration risks, and defining standards other engineers build against—measured in cluster reliability, time-to-first-GPU on new capacity, and quality at scale. You will mentor senior and junior engineers, deepen the team's expertise in virtualization, DC networking, and GPU infrastructure, and help raise the hiring bar. You will create testing frameworks, tools, and developer documentation that make systems robust and usable by other teams, and shape the core, open-source Together AI platform.
Success requires deep technical expertise and excellent communication. You must have expert software development fundamentals, deep systems knowledge and troubleshooting instincts, and the leadership and diplomacy skills to align teams that don't report to you. Much of this work starts ambiguous; you will define scope and drive it to production.
**Requirements:**
- 7+ years of professional software development experience with expert-level proficiency in at least one backend language (Golang desired); ability to write high-performance, well-tested, production-quality code
- Track record of owning the architecture of large distributed systems from blank page to production at scale, including judgment calls that could not be reversed cheaply
- Deep experience building and operating globally distributed, high-performance microservice architectures across one or more cloud providers (AWS, Azure, GCP)
- Expert systems knowledge across compute, networking, and storage—including concurrency, memory management, performant I/O, and scale at a global level
- Demonstrated technical leadership beyond your own commits: mentoring senior engineers, leading design reviews, and driving alignment across teams that do not report to you
- Excellent communication and diplomacy skills—able to write design docs that settle arguments and work effectively with technical and non-technical stakeholders
- Experience building and operating reliable, customer-facing production systems at scale, and owning infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD)
**Preferred Qualifications:**
Deep Kubernetes internals experience (implementing Kubernetes operators, device/storage/network plugins, custom schedulers, or patches); deep experience with VMs/hypervisors (QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV); deep experience with DC networking tech (VLAN, VXLAN, VPN, VPC, OVS/OVN); experience with Cluster API; experience with high-performance compute, networking, and/or storage; experience virtualizing GPUs and/or InfiniBand; experience building IaaS or PaaS systems at scale; experience with DPUs/SmartNICs; GPU programming, NCCL, CUDA knowledge.
About Together AI
AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.