SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Hippocratic AI is building the leading generative AI platform for healthcare, with a focus on safe, autonomous clinical conversations. The company has developed proprietary LLMs (Polaris constellation) achieving over 99.9% accuracy and recently raised $126M Series C at $3.5B valuation.
As Staff Applied Scientist for Reinforcement Learning, you will own the end-to-end RL and On-Policy Distillation (OPD) post-training pipeline. This role is critical to transforming raw model capability into reliable, safe clinical behavior—where stakes are exceptionally high in healthcare.
Key responsibilities:
- Design and implement advanced RL and OPD post-training methods (RLHF, RLVR, OPD variants)
- Build and evaluate reward models, verifiers, and LLM-as-judge evaluation pipelines
- Develop conversational AI environments and simulations for healthcare-specific RL training using synthetic data
- Automate post-training loops with agents for continuous research and improvement
- Run rigorous experiments to isolate and understand drivers of post-training gains
- Collaborate cross-functionally with research, engineering, and clinical teams to ensure models meet safety and clinical reasoning standards
Required qualifications:
- MS or PhD in Computer Science or related field
- 5+ years experience in NLP, LLM training, or reinforcement learning
- 2+ years hands-on experience in RL for LLM post-training
- Proven experience with large-scale LLM training (50B+ parameters, multi-node distributed systems)
- Strong Python and PyTorch proficiency
- Demonstrated expertise with RLHF, RLVR, LLM-as-judge, or similar post-training methods
Nice-to-have:
- Publications at top-tier venues (NeurIPS, ICML, ICLR, ACL, EMNLP)
- Healthcare domain experience or familiarity with clinical AI applications
The role is based in Palo Alto with an expectation of five days per week in-office collaboration. You'll work alongside physicians, hospital leaders, AI researchers, and engineers from institutions like Stanford, Johns Hopkins, Google, Meta, Microsoft, and NVIDIA.