SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior AI Engineer - Core

Hilbert - San Francisco, CA, United States - In-office - posted 2026-09-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Hilbert is building a reasoning engine and demand intelligence platform that orchestrates multi-step inference over enterprise data to turn months-long decision cycles into minutes. The platform is fully agentic by design and powers growth operations for Fortune 500 enterprises and brands like FreshDirect, Blank Street, and Levain Bakery. You will own core pieces of the AI stack end-to-end—from prototype to production pipeline to product. This role focuses initially on building the evaluation and testing layer for production agents, then expands into retrieval, orchestration, and execution across the full AI stack. Key responsibilities: - Design and own the evaluation layer for agents: harnesses, metrics, golden datasets, regression gates, and human-in-the-loop review systems - Architect and implement agent workflows using LangChain, LangGraph, or equivalent frameworks; manage state memory, routing, tool registries, and recovery paths - Own systems from experimentation through production: implement tracing, monitoring, latency optimization, cost-per-task budgeting, and on-call support - Diagnose production failures (hallucination, tool misuse, retrieval misses, silent degradation) and convert each into durable fixes and test cases - Set technical standards for agent work: review designs, define patterns, and raise the bar across the team - Collaborate with the founding team and cross-functional partners to communicate tradeoffs, progress, and technical decisions - Make pragmatic engineering decisions under ambiguity; ship, learn, and iterate at startup speed Day-one challenges include: intelligent retrieval across heterogeneous approaches (combining RAG, graph-based retrieval, and other methods); building agentic workflows robust enough to handle edge cases and unexpected situations; creating systematic, reproducible evaluations that predict real-world performance; and integrating agents with external platforms for real-world execution while maintaining human-in-the-loop checkpoints. REQUIREMENTS: Must-haves: - 6+ years of production software engineering (APIs, services, data infrastructure); you've owned code others depended on with tests, CI/CD, and on-call responsibility - 2+ years building LLM or agent systems shipped to production with real user adoption (not internal demos or prototypes); ability to discuss production failures and resolutions - Hands-on experience with LangChain, LangGraph, or equivalent agent/orchestration frameworks; you've built with them, hit their limits, and worked around them - Clear communication and conviction; ability to explain technical decisions to non-technical founders and debate architecture tradeoffs with senior engineers - Ownership mindset; you see what needs building, raise your hand, and ship it without waiting for tickets - Thrive in ambiguity; energized by evolving requirements and fast-moving AI products - Move at startup speed; know which decisions deserve a day of thought versus an hour Strong pluses: - Deep RAG expertise: hybrid and graph retrieval, chunking and embedding strategy, ranking, grounding - Observability for LLM systems (Langfuse, OpenTelemetry, or equivalent) and cost/latency optimization - MCP, tool-calling frameworks, structured output, and constrained decoding - Experience at early-stage startups or high-growth environments wearing multiple hats

Similar roles