SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Instabase is building AI-powered infrastructure to help organizations extract insights from unstructured data at scale. The company serves enterprise customers and is backed by top-tier VCs including Greylock and Andreessen Horowitz.
You will join the distributed systems and AI agents team, which tackles some of the most challenging technical problems in scaling agentic AI capabilities. The team works on three interconnected areas: distributed systems for large-scale unstructured-data extraction workloads, reliable backend infrastructure for agentic execution and long-running AI workflows, and R&D into advances in LLMs and AI.
In this role, you will take new ideas from 0-to-1 quickly, then evolve their architecture to scale by orders of magnitude—from early prototypes to production systems handling workloads that are 10×, 1,000×, or even 1,000,000× larger. You'll build scalable distributed systems for large-scale unstructured-data extraction and agentic workloads, develop reliable execution infrastructure for scheduling, retries, timeouts, cancellation, rate limits, and failure recovery, and build agent runtimes that coordinate model calls, tool execution, durable state, multi-step loops, and human approvals. You'll also evaluate emerging models and AI techniques through benchmarks and experiments, and apply them to new platform and customer problems. You will own systems from design through production, including testing, observability, and pragmatic trade-offs across reliability, performance, security, and cost.
You bring 5+ years of software engineering experience building and operating production backend systems, with the ability and experience to serve as a technical lead—setting direction, leading teams through complex projects, and mentoring and coaching engineers while remaining hands-on. You have strong fundamentals in distributed systems, concurrency, APIs, data modeling, and failure handling. You have experience with agentic execution systems, workflow orchestration, event processing, scheduling, or high-volume services. You are proficient in a backend language such as Python, Go, Java, Kotlin, Rust, or C++. You have experience with queues, databases, caching, containers, cloud infrastructure, and production observability. You have experience building or integrating ML, generative-AI, or LLM-based systems, including tool calling and agent workflows. You have a strong ownership mindset across design, deployment, monitoring, incident response, and continuous improvement.