SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Hilbert is building a reasoning engine and demand intelligence platform that orchestrates multi-step inference over enterprise data, turning months-long decision cycles into minutes. The platform powers growth operations for Fortune 500 enterprises and brands like FreshDirect, Blank Street, and Levain Bakery.
You will own core pieces of the AI stack end-to-end—from prototype to production pipeline to product. This role focuses initially on building the evaluation and testing layer for production agents, then expands into retrieval, orchestration, and execution across the full AI stack.
Key responsibilities:
- Own the evaluation layer for agents: design harnesses, define metrics and golden datasets, build regression gates, and implement human-in-the-loop review systems
- Architect and implement agent workflows using LangChain, LangGraph, or equivalent frameworks; manage state, routing, tool registries, and recovery paths
- Operate systems from experimentation through production: implement tracing, monitoring, latency optimization, cost-per-task budgeting, and on-call support
- Diagnose and fix production failures (hallucination, tool misuse, retrieval misses, silent degradation) and convert each into durable fixes and test cases
- Set technical standards for agent work: review designs, define patterns, and raise the bar across the team
- Collaborate with the founding team and cross-functional partners on technical tradeoffs and decisions
- Make pragmatic engineering decisions under ambiguity and ship fast
Key problems you'll solve:
- Intelligent retrieval across heterogeneous approaches: combine RAG, graph-based retrieval, and other methods into a unified strategy
- Robust agentic workflows that handle edge cases, missing data, and unexpected situations through reasoning and human escalation
- Systematic, reproducible evaluation frameworks that predict real-world performance
- Execution and real-world integration where agents take action with human-in-the-loop checkpoints
Requirements:
- 4+ years of production software engineering (APIs, services, data infrastructure); you've owned code others depended on with tests, CI/CD, and on-call responsibility
- 2+ years building LLM or agent systems shipped to production with real users; not internal demos or prototypes
- Hands-on experience with LangChain, LangGraph, or equivalent agent/orchestration frameworks; you've hit their limits and worked around them
- Strong communication skills: explain technical decisions to non-technical founders and debate architecture tradeoffs with senior engineers
- Ownership mindset: you see what needs building, raise your hand, and ship it without waiting for tickets
- Thrive in ambiguity; energized by evolving requirements and fast-moving AI products
- Move at startup speed; know which decisions deserve a day of thought versus an hour
Strong pluses:
- Deep RAG experience: hybrid and graph retrieval, chunking and embedding strategy, ranking, grounding
- Observability for LLM systems (Langfuse, OpenTelemetry, or equivalent) and cost/latency optimization
- MCP, tool-calling frameworks, structured output, and constrained decoding
- Experience at early-stage startups or high-growth environments wearing multiple hats
Timezone requirement: at least 5 hours overlap with PST timezone (7am-5pm).