SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Eucalyptus (now part of Hims & Hers) is a digital healthcare company focused on obesity and weight management. The company operates Juniper, one of the world's largest weight-management programs combining GLP-1 medication with personalized nutrition, movement support, and clinician-led care. Eucalyptus operates across five markets (Australia, UK, Germany, Japan, Canada) and has received NICE endorsement for NHS services.
In this role, you will own the development of a standardized evaluation framework for AI and human conversations across all patient-facing support workstreams. This includes clinical interactions (consultations, medical support), health coach conversations, AI assistant conversations, patient experience support, and health insights. You will build services that monitor patient-facing conversations and automatically flag risky content for review and escalation. You'll develop evaluation datasets, metrics, and thresholds in collaboration with clinicians, coaches, and support leads. You'll own the technical integration between AI assistants and internal patient data models, ensuring assistants respond with accurate, well-grounded patient context. You'll operate systems in production with monitoring, alerting, and iteration as models and prompts evolve. You'll also own and automate guardrail updates based on internal and patient feedback, and share results with teams to help them improve.
This is positioned as a rare junior role with real ownership of a safety-critical system from day one, supported by experienced teams. You'll report to the Head of Data and work closely with ML Engineering, Data Engineering, and Data Science teams.
REQUIREMENTS:
- Degree in computer science, engineering, data science, or related field (or equivalent practical experience)
- Strong Python skills and fundamentals to build reliable services, not just notebooks
- Hands-on experience with LLMs: prompt engineering, and ideally exposure to eval frameworks or LLM-as-judge approaches (coursework, internships, or personal projects count)
- Safety-first mindset; care about what AI systems say to real patients and attention to edge cases
- Clear communication skills; ability to explain eval results to clinicians and engineers
- Comfortable with ambiguity; willingness to help define a system that doesn't yet exist
NICE TO HAVE:
- Experience with eval tooling (promptfoo, Ragas, LangSmith, or in-house harnesses)
- Exposure to healthcare, regulated industries, or other high-stakes AI applications
- Familiarity with GCP, SQL, or modern data stacks