SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Together AI is seeking a Staff Software Engineer to design and build the infrastructure provisioning systems that power their AI inference and training clusters. This role owns the complete lifecycle of physical hardware—from bare-metal discovery through GPU driver setup, health validation, and decommissioning—modeling it as explicit, versioned state machines that can be reconciled automatically.
You will build a declarative, self-service API that allows the Research and Inference teams to provision, scale, and tear down clusters with a single API call, eliminating manual runbook execution. The platform is manifest-driven: teams declare desired cluster state (topology, software stack, shape), and your systems continuously reconcile reality to that manifest, similar to how Kubernetes controllers work.
Key responsibilities include designing the provisioning state machine with explicit transitions for each lifecycle stage; building the self-service control plane and APIs; implementing self-healing capabilities to detect, drain, and repair failed nodes automatically; ensuring reliability through idempotency, retries, rollback, and drift detection; and partnering with the ML platform team to encode their cluster requirements as first-class abstractions.
This is infrastructure-as-software: you will write typed, tested, versioned code deployed through CI/CD—not Ansible playbooks or manual processes. You own both the software delivery and production operation.
Required: strong software engineering background in Go, Python, Rust, or similar; experience with durable workflow orchestration (Temporal, Cadence, or equivalent); experience building control planes or orchestration systems that model and reconcile state; event-driven systems design (Kafka, NATS, SQS); and a product mindset with experience building internal platforms for other engineering teams.
Nice-to-have: bare-metal provisioning (PXE, Redfish, IPMI), networking fundamentals (VLANs, BGP), GPU/accelerator infrastructure, GPU cluster software stacks (NCCL, CUDA, InfiniBand), hyperscaler or datacenter-scale experience, and systems programming in Rust or Go.
About Together AI
AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.