SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer — Infra Agent Systems UK

Together AI - Remote - Remote - posted 2026-09-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Together AI operates one of the world's largest GPU fleets and is building the Infra Agent Systems team to develop production AI agents that diagnose hardware failures, investigate incidents, and automate operational workflows across that infrastructure. In this role, you will work across two interconnected areas: **Infrastructure Agent Systems**: Design and build production AI agents that help operate the GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are deployed daily through APIs, CLIs, dashboards, and Slack integrations used by infrastructure and datacenter teams. **Core Agent Platform**: Build the foundational platform powering these agents, including knowledge graphs, search and retrieval systems, orchestration frameworks, evaluation systems, and developer tooling that enables agents to reason, act, and continuously improve. Key responsibilities include designing distributed services and orchestration frameworks, developing fleet intelligence systems that combine telemetry and operational knowledge, integrating with observability and incident management systems, owning services end-to-end from architecture through production operations, and improving agent performance through evaluations and production feedback loops. You'll work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. The role requires 5+ years building production backend or distributed systems, strong systems design skills, and depth in at least one of: AI agent systems/orchestration, knowledge graphs, or search/retrieval/RAG systems. Strong backend engineering experience with API design, service boundaries, and data modeling is essential. Experience with Kubernetes, GitOps (ArgoCD), infrastructure-as-code, and cloud platforms is required. Proficiency across Go, TypeScript, Python, or Rust is expected. Plus experience includes GPU infrastructure, datacenters, bare-metal systems, graph databases, event-driven systems (NATS/Kafka), observability platforms (Prometheus/Grafana), and building evaluation frameworks for LLM-powered systems.

About Together AI

AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.

Similar roles