SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer — Infra Agent Systems

Together AI - San Francisco, CA, United States - Hybrid - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 250,000 - 300,000 / annual

Together AI operates one of the world's largest GPU fleets and is building the software systems that power and automate that infrastructure. The Infra Agent Systems team is hiring a Senior Software Engineer to work on two interconnected areas: **Infrastructure Agent Systems**: Build production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. These agents are used daily by infrastructure and datacenter teams through APIs, CLIs, dashboards, and Slack integrations. **Core Agent Platform**: Develop the platform powering these agents, including knowledge graphs, retrieval systems, orchestration frameworks, evaluation tools, and developer tooling that enables agents to reason, act, and continuously improve. This role sits at the intersection of AI agents, distributed systems, infrastructure, and automation. You'll tackle two hard problems simultaneously: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You'll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure. **Key Responsibilities**: - Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world's largest GPU fleets - Build distributed services, orchestration frameworks, knowledge graphs, and retrieval systems powering infrastructure agents - Develop fleet intelligence systems combining telemetry, infrastructure state, operational knowledge, and historical incidents - Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs - Own services end-to-end: architecture, implementation, testing, deployment, observability, and production operations - Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops - Turn what agents learn in production into reliable, reviewed software and automation **Requirements**: - 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms - Strong systems design skills and experience owning significant systems from design through production - Depth in at least one of: AI agent systems (orchestration, tool use, evaluation, grounding), knowledge graphs or graph data modeling, or search/retrieval/ranking/RAG/semantic search systems - Strong backend engineering experience including API design, service boundaries, data modeling, and integrations across complex systems - Experience with Kubernetes, GitOps (such as ArgoCD), infrastructure-as-code, and cloud platforms - Comfortable working across languages such as Go, TypeScript, Python, or Rust **Nice-to-Have**: - GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers - Graph databases - Event-driven systems and messaging platforms (NATS, Kafka) - Observability platforms (Prometheus, Grafana) - Building evaluation frameworks or improving quality and reliability of LLM-powered systems

About Together AI

AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.

Similar roles