SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Arena Intelligence is building the platform for evaluating how AI models perform in real-world scenarios. Founded by researchers from UC Berkeley's SkyLab, the company enables tens of millions of users monthly to evaluate frontier AI systems. The platform powers transparent, human-centered evaluations used by leading AI labs, enterprises, and independent researchers.
You'll join the infrastructure team as an early member, building the core systems that enable Arena's online evaluation platform to operate at scale. This is a hands-on individual contributor role focused on zero-to-one and scale-up infrastructure work.
Key responsibilities include:
- Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas from the ground up
- Solve complex streaming problems: handle SSE/streaming responses across heterogeneous AI providers, including partial failure recovery, mid-stream fallback, and consistent response normalization
- Build enterprise-grade infrastructure including rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance as the company scales beyond individual developers
- Instrument infrastructure with deep observability: distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards
- Integrate with Arena's core evaluation platform and collaborate with the research team to turn novel ideas into full-featured products
- Contribute to backend systems across the Leaderboards and Evals platforms, helping unify public and private data architectures
The role involves working closely with researchers, engineers, and product leadership in a fast-moving startup environment. The AI gateway is currently in private launch with individual developers; enterprise use cases and new infrastructure features are shipping in the near term.
REQUIREMENTS:
- ~5+ years of backend engineering experience with meaningful time on distributed systems, infrastructure, or developer-facing platforms
- Strong proficiency in Go (primary backend language — must-have)
- Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and understanding of streaming, token management, rate limits, and model-specific quirks
- Product-oriented mindset: think about developer experience of APIs, ask "why" before "how"
- Comfort with ambiguity in a startup environment where scope is fluid and context shifts
NICE TO HAVE:
- Cloud infrastructure experience (AWS, GCP, Azure), Kubernetes, Terraform
- Database systems experience (Postgres, Redis)
- API gateway, proxy, or developer tools experience (Bifrost, Kong, Envoy, Tyk)
- AI/ML infrastructure, model serving, inference, or evaluation framework background
- Enterprise-ready features experience (SSO, RBAC, audit logs, multi-tenancy)
- Billing infrastructure experience (Stripe, Metronome, Orb)
- Familiarity with modern AI infra stack (vLLM, LiteLLM, LangChain)