SlipstreamJobsFresh Startup & VC-Backed Jobs

AI Engineer - Core

Hilbert - San Francisco, CA, United States - In-office - posted 2026-09-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Hilbert is building a reasoning engine and demand intelligence platform that orchestrates multi-step inference over enterprise data to turn months-long decision cycles into minutes. The platform is fully agentic by design and powers growth operations for Fortune 500 enterprises and brands like FreshDirect, Blank Street, and Levain Bakery. You will own core pieces of the AI stack end-to-end—from prototype to production pipeline to product. This is a systems-level role where you design agent architectures, build evaluation systems, and make hard tradeoffs between accuracy, latency, and cost in production environments with real enterprise customers. Initially, you'll own the evaluation and testing layer for Hilbert's production agents. Before expanding agent capabilities, you'll design the eval harness, define correctness metrics for multi-step agent trajectories, build regression gates, and convert production failures into durable test cases. From there, your scope expands into retrieval, orchestration, and execution across the full AI stack. Key responsibilities: - Design and own the evaluation layer: harnesses, metrics, golden datasets, regression gates, and human-in-the-loop review systems - Architect and implement agent workflows using LangChain, LangGraph, or equivalent frameworks; manage state, routing, tool registries, and recovery paths - Own systems from experimentation through production: implement tracing, monitoring, latency optimization, cost-per-task budgeting, and on-call support - Diagnose production failures (hallucination, tool misuse, retrieval misses, silent degradation) and convert each into a durable fix and test case - Set technical standards for agent work: review designs, define patterns, and raise the bar across the team - Collaborate with the founding team and cross-functional partners on technical tradeoffs and decisions - Make pragmatic engineering decisions under ambiguity; ship, learn, iterate at startup speed Key problems you'll solve: - Intelligent retrieval across heterogeneous approaches: combining RAG, graph-based retrieval, and other methods into a unified strategy that fetches the right information at the right moment - Building agentic workflows robust enough to handle edge cases, missing data, and unexpected situations; reasoning through problems and escalating to humans when necessary - Systematic, reproducible evaluation that predicts real-world performance beyond subjective assessment - Execution and real-world integration: agents that take action, integrate with external platforms, and execute workflows with human-in-the-loop checkpoints Requirements: - 4+ years of production software engineering (APIs, services, data infrastructure); you've owned code that others depended on with tests, CI/CD, and on-call responsibility - 2+ years building LLM or agent systems shipped to production with real users; not internal demos or prototypes. You can discuss production failures and how you resolved them - Hands-on experience with LangChain, LangGraph, or equivalent agent/orchestration frameworks; you've built with them, hit their limits, and worked around them - Clear communication: you can explain technical decisions to non-technical founders and debate architecture tradeoffs with senior engineers - Ownership mindset: you don't wait for tickets; you identify what needs building and ship it - Comfort with ambiguity: AI products evolve rapidly and you're energized by figuring things out - Startup speed: you move fast, make decisions without permission, and know which decisions deserve a day of thought versus an hour Strong pluses: - Deep RAG experience: hybrid and graph retrieval, chunking and embedding strategy, ranking, grounding - Observability for LLM systems (Langfuse, OpenTelemetry, or equivalent) and cost/latency optimization - MCP, tool-calling frameworks, structured output, and constrained decoding - Experience at early-stage startups or high-growth environments wearing multiple hats Location: San Francisco with occasional travel for team meetings, offsites, and customer engagements.

Similar roles