SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, AI Evaluation

Nuna - San Francisco, CA, USA - In-office - posted 2026-08-13

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Nuna is building an AI health coach to help the 130 million Americans managing chronic conditions. Unlike traditional healthcare's 15-minute clinic visits, Nuna's coach is available 24/7, uses motivational interviewing, and helps patients design experiments that fit their real lives. You'll own the evaluation system end-to-end for Nuna's AI agents. This is a net-new, build-first role where you design and implement the infrastructure, datasets, judges, and release gates that determine whether the coach is safe and effective before deployment. You won't execute tests someone else designed—you'll write the code, own the system, and make day-to-day calls on standards and trade-offs with significant autonomy. Key responsibilities include: building testing harnesses and evaluation infrastructure for agentic products; owning evals architecture and content with support from data science and clinical partners; ensuring every deployment passes testing before reaching patients; building ground truth, judges, and metrics while validating evaluation reliability; creating functional tooling for clinicians and designers to author and review scenarios without engineering overhead; and closing the loop from evaluation results to model and prompt refinement. You'll work within a small, interdisciplinary team of engineers, data scientists, designers, product managers, and clinicians. Your evaluation signals directly inform what ships to patients, making this a high-impact role in a regulated healthcare context. Required: significant production systems experience; deep expertise in AI evaluation (LLM-as-judge, red-teaming, synthetic scenario generation, multi-turn/agentic evaluation); understanding of how evals themselves fail; daily AI usage and tool-building; statistics and experimental design fluency to partner with data scientists; ability to design workflows and build functional UIs for non-engineers; genuine interest in healthcare improvement with judgment to distinguish launch-blocking issues from nice-to-haves. Bonus: healthcare or regulated domain experience; hands-on eval tooling ecosystem familiarity (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo); red-teaming or AI safety experience; automated eval-driven optimization; early-stage environment experience.

Similar roles