SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Backend Engineer, Vision

Sarvam AI - Bengaluru, India - In-office - posted 2026-08-18

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sarvam AI is building India's sovereign AI platform, developing full-stack AI infrastructure across research, models, and applications. The company partners with leading Indian enterprises including Tata Capital, SBI Life, CRED, and IDFC, and is backed by Lightspeed, Peak XV, and Khosla Ventures. You will own the end-to-end architecture of the serving harness for Sarvam's vision-language models—the production system that delivers frontier-grade document extraction quality from 3B and 30B in-house models at national scale. This is fundamentally an engineering problem: closing the gap between small sovereign models and expensive frontier hosted models (Gemini Flash) through intelligent system design. The role involves architecting and operating systems that process real Indian documents at population scale: PAN, Aadhaar, bank statements, GST filings, insurance reports, contracts, and legal documents across multiple languages and scan qualities. You will own the hard trade-off surface directly—accuracy versus latency versus cost per page—and control the levers: multi-pass inference, model routing and cascades, self-consistency verification, confidence-driven escalation, batching, caching, and GPU utilization. Key responsibilities include: designing the accuracy harness with multi-pass extraction, ensembling, cross-verification, and confidence calibration; architecting durable, resumable document workflows in Temporal with fan-out, partial failure recovery, and exactly-once semantics; owning the inference serving layer with batching strategy, GPU pool management, autoscaling, and multi-model routing; driving latency, throughput, and unit economics down through profiling and measurement; building observability infrastructure with distributed tracing, per-stage metrics, SLOs, and alerting; designing for multi-tenancy, tenant isolation, and fair scheduling; supporting on-prem and constrained deployments; and mentoring SDE 1–2 engineers through design review and technical direction. You will work with a high-talent-density team of researchers, engineers, and builders moving fast on problems with real population-scale impact. The tech stack includes Go, Python, Temporal, Kubernetes, PostgreSQL, Redis, and OpenTelemetry-based observability. Required: 5–6+ years backend engineering with meaningful time operating high-throughput production systems on-call; deep proficiency in Go and/or Python; strong distributed systems design (queues, orchestration, idempotency, backpressure, retry semantics, consistency, graceful degradation); production experience with Temporal or equivalent durable execution engines; Kubernetes in production including autoscaling, resource management, and GPU workload scheduling; demonstrable experience serving ML/LLM inference in production (batching, caching, model versioning, A/B rollout, latency budgeting); rigour about observability and reliability (SLO design, incident response, shipping fixes); and hard-won cost intuition. Bonus: direct experience with OCR, IDP, or document AI systems; GPU inference stacks (vLLM, TensorRT-LLM, Triton, SGLang, Ray Serve); ML evaluation infrastructure; BFSI/healthcare/public-sector compliance experience; on-prem or air-gapped deployment experience.

Similar roles