SlipstreamJobsFresh Startup & VC-Backed Jobs

NLP / LLM Data Scientist

Dandelion Health Inc - Remote - Remote - posted 2026-09-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Dandelion Health is building the world's largest AI training and clinical development platform, partnering with major U.S. health systems to safely and ethically make de-identified patient data available to AI developers and pharma companies. The company was founded in 2020 by experts in health tech, hospital systems, academia, and clinical AI, and currently works with Sharp HealthCare, Sanford Health, and Texas Health Resources, with additional health systems joining soon. The company has access to clinical data dating back to July 2016, representing over 10 million patients. This includes structured EMR data, unstructured clinical notes and radiology reports, medical images (DICOM, pathology), video, waveforms, and continuous streaming monitoring data. As an NLP / LLM Data Scientist, you will be a healthcare data scientist responsible for leveraging and building upon large language models and other ML-based approaches for meaningful data abstraction from unstructured and structured healthcare data. You will join a team of data scientists who own the creation and maintenance of AI-ready datasets for clients. Your work will span curating datasets by identifying patient subpopulations or disease cohorts, developing methodology to abstract information from multimodal healthcare data sources to support patient phenotyping, and pooling data to enable rapid exploratory AI/ML analyses, model experimentation, and model validation. Key responsibilities include: developing NLP, LLM, and other ML-based pipelines to abstract relevant labels from text-based healthcare data; querying complex source systems across EMRs, semi-structured reports, and free-text clinical notes to identify key data elements and create high-quality datasets; owning data extraction, wrangling, labeling, and QC tasks; staying current on applied NLP and generative AI methods; supporting design, testing, validation, and analysis of multimodal data structures; developing HIPAA-compliant code and documentation; identifying and resolving problems; ensuring data accuracy and integrity; supporting the technical product team; and communicating findings clearly to audiences with varying technical and clinical backgrounds. You will report to the Data Science Manager, under the Head of Data. The work is fast-paced and iterative. The company values experimentation, diversity of perspectives, and learning from data. Occasional travel for in-person company working days on a quarterly basis is expected. QUALIFICATIONS: - Advanced degree in a quantitative field (Data Science, Biomedical Informatics, Computer Science, Biostatistics) OR B.S. with at least 5 years of professional experience - At least 2 years of data science and machine learning experience, including building pipelines to extract and curate unstructured and semi-structured data using advanced ML and AI techniques; prior clinical/healthcare data experience is a strong bonus - Fluency in Python and SQL, including ML/NLP libraries (PyTorch, TensorFlow, HuggingFace, etc.) - Familiarity with modern applied LLM techniques on real-world data - Strong technical writing, editing, and communication skills with a collaborative mindset - Excellent organizational skills with ability to manage multiple projects and meet deadlines - Startup experience is a plus Desired technical skills (not all required): Git and version control, encryption methods, EDW/database querying experience, knowledge of EMR systems (Epic, Cerner, Allscripts), medical ontology experience, DICOM or imaging modality experience, AWS, peer-reviewed publication experience.

Similar roles