SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Hippocratic AI is a generative AI company focused on healthcare, building the first healthcare-only, safety-focused LLM system capable of safe, autonomous clinical conversations with patients. The company has achieved over 99.9% accuracy in its Polaris constellation of models and recently raised $126M in Series C funding at a $3.5B valuation.
In this Senior Applied Scientist role, you will own the end-to-end Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline. Your work will directly improve the models' clinical reasoning, safety, and alignment—critical for a healthcare AI system deployed to millions of patients across diverse clinical use cases.
Key responsibilities include:
- Designing RL and OPD post-training methods such as RLHF, RLVR, and OPD
- Building and evaluating reward models, verifiers, and LLM-as-judge pipelines
- Developing conversational AI environments and simulations for healthcare RL training using synthetic data
- Automating post-training loops with agents for auto-research
- Running rigorous experiments to understand drivers of post-training gains
- Collaborating with research, engineering, and clinical teams
Required qualifications:
- MS or PhD in Computer Science or relevant field
- 3+ years of experience in NLP, LLM training, or RL
- 1+ years of experience in RL for LLM post-training
- Experience with large-scale LLM training (50B+ parameters, multi-node)
- Strong Python and PyTorch coding skills
- Experience with RLHF, RLVR, LLM-as-judge, or similar LLM post-training methods
The role is based in the Palo Alto office with an expectation of five days per week on-site to support collaboration and team culture. You'll work alongside physicians, hospital leaders, AI pioneers, and researchers from institutions including Stanford, Google, Meta, Microsoft, and NVIDIA.