SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sapiom is building the end-to-end platform for shipping and scaling agentic products. The company unifies compute, memory, identity, spend controls, storage, queues, and monitoring into a single operating system for autonomous agents. Founded by the former head of payments engineering at Shopify, Sapiom raised a $35M Series A led by Dragonfly (total $50M), with backing from Accel, Menlo Ventures, and Anthropic.
The core problem: agents can think but can't act reliably or economically at scale. Teams build demos quickly but struggle in production—systems break unpredictably, costs spiral, and observability is poor. Sapiom closes this gap by providing infrastructure that makes agent execution economical, reliable, and controllable.
As a Software Engineer on the Agent Infrastructure team, you'll work on several interconnected systems: routing and capacity management (placing model calls against latency, cost, and quality targets in real time); metering and billing correctness (handling 270M+ consumption events with ledger reconciliation); reliability at sustained scale (the company has experienced 10–20x growth compounding week-over-week); the control plane (credential-scoped permissions, workflow state limits, gateway hardening); and agent execution (sandboxed runs, state, memory, tools, and third-party integrations at tens of thousands of runs daily).
The organization is flat and early-stage. There's no fixed lane—you might work on routing one month and metering the next. You'll have direct surface area with architects and founders, meaning your design decisions get questioned by experienced engineers who've made these mistakes before. This accelerates learning but requires strong fundamentals and the ability to move fluidly between backend, infrastructure, and tooling.
You should have shipped something in production that you were responsible for when it broke. You're comfortable taking problems from "this should exist" to shipped with minimal structure. You make reasonable calls with incomplete information and revise them as you learn. You close gaps fast—learning what you're missing rather than routing around it. Experience with distributed systems, queues, storage, observability, third-party API integrations, LLM inference, or anything you've built and operated at real scale is valuable. Using AI tools as a multiplier in your own workflow is a plus.