SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI - London, United Kingdom - In-office - posted 2026-08-20

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Together AI is seeking a Staff Software Engineer to design and build the infrastructure provisioning platform that powers their AI inference and training clusters. This role owns the software state machines that automate the full lifecycle of physical hardware—from discovery and GPU driver installation through health validation to decommissioning—eliminating manual provisioning work entirely. You will architect declarative APIs and control planes that allow the Research and Inference teams to request, scale, and tear down inference clusters with a single API call, with no human intervention required. The platform is manifest-driven: teams declare desired cluster state (shape, topology, software stack), and your systems continuously reconcile reality to that manifest, similar to how Kubernetes controllers work. Key responsibilities include building the provisioning state machine with explicit, versioned states and transitions; designing self-service APIs for cluster lifecycle management; implementing self-healing mechanisms that detect degraded nodes, drain them safely, and trigger repairs automatically; ensuring reliability through idempotency, retries, rollback, and drift detection; and partnering with the ML platform team to encode their cluster topology and scheduling requirements as first-class abstractions. You will write production-grade infrastructure code—strongly typed, tested, versioned, and deployed through CI/CD—treating infrastructure as software rather than as runbooks or Ansible playbooks. You own both the software delivery and production operation of what you build. Required experience includes strong software engineering in Go, Python, Rust, or similar languages; hands-on work with durable workflow orchestration tools (Temporal, Cadence, or equivalent); building control planes or orchestration systems that model and reconcile state; designing event-driven systems around message queues or pub/sub (Kafka, NATS, SQS); and a product mindset with experience building internal platforms consumed by other engineering teams. Nice-to-have skills include bare-metal provisioning (PXE, Redfish, IPMI, BMC), networking fundamentals (VLANs, BGP, fabric design), GPU/accelerator infrastructure, GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE), prior hyperscaler or datacenter-scale infrastructure experience, and systems programming in Rust or Go.

About Together AI

AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.

Similar roles