SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Together AI is seeking a Staff Software Engineer to design and build the infrastructure provisioning platform that powers their AI inference clusters. This role owns the software state machines that manage the full lifecycle of physical hardware—from bare-metal discovery through GPU driver/CUDA stack setup, health validation, and decommissioning—transforming racks of GPUs into fully operational inference clusters without manual intervention.
You will architect a declarative, manifest-driven control plane that allows the Research and Inference team to provision, scale, and tear down clusters via a single API call. The platform will use durable workflow orchestration (Temporal, Cadence, or equivalent) to execute long-lived, failure-resilient provisioning workflows. You'll design state machines and reconciliation engines similar to Kubernetes controllers, ensuring idempotency, retries, rollback, and continuous drift detection.
Key responsibilities include: building the provisioning state machine with explicit versioned states and transitions; designing self-service declarative APIs; automating self-healing (node failure detection, safe draining, repair triggering, and reintroduction); ensuring reliability through production-grade software practices (typing, testing, versioning, CI/CD); and partnering with the ML platform team to encode cluster topology and scheduling constraints as first-class abstractions.
This is a product-minded role where you build internal platforms consumed by other engineering teams. You own both the software delivery and production operation. The platform is engineered as real software—not Ansible playbooks—with strong typing, automated tests, code review, and CI/CD pipelines.
Required: strong software engineering background in Go, Python, Rust, or similar; experience with durable workflow orchestration tools; experience building control planes or orchestration systems that model and reconcile state; event-driven systems design (Kafka, NATS, SQS); product mindset with internal platform/API experience.
Nice-to-have: bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC), networking fundamentals (VLANs, BGP), GPU/accelerator infrastructure, GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE), hyperscaler or datacenter-scale infrastructure experience, systems programming in Rust or Go.
About Together AI
AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.