SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Together AI operates one of the world's largest GPU fleets and is building the Infra Agent Systems team to develop production AI agents that diagnose hardware failures, investigate incidents, and automate operational workflows across that infrastructure.
In this role, you'll work across two interconnected areas:
**Infrastructure Agent Systems**: Design and build production AI agents that help operate the GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used daily by infrastructure and datacenter teams through APIs, CLIs, dashboards, and Slack integrations.
**Core Agent Platform**: Build the foundational platform powering these agents, including knowledge graphs, search and retrieval systems, orchestration frameworks, evaluation tools, and developer tooling that enables agents to reason, act, and continuously improve.
Key responsibilities include designing distributed services and orchestration frameworks, developing fleet intelligence systems that combine telemetry and operational knowledge, integrating with observability and incident management systems, owning services end-to-end from architecture through production operations, and improving agent performance through evaluations and feedback loops.
You'll work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. This is an opportunity to build foundational systems from the ground up and help define how self-improving AI agents operate real-world AI infrastructure at massive scale.
Required: 5+ years building production backend or distributed systems; strong systems design skills; depth in AI agent systems, knowledge graphs, or search/retrieval systems; strong backend engineering experience; familiarity with Kubernetes, GitOps, infrastructure-as-code, and cloud platforms; comfort with Go, TypeScript, Python, or Rust.
Plus: GPU infrastructure experience, graph databases, event-driven systems (NATS/Kafka), observability platforms, or LLM evaluation frameworks.
About Together AI
AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.