SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Data Scientist

Mainstay - Remote - Remote

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 150,000 - 165,000 / annual

Mainstay, a division of Lemnis (a public charity focused on expanding learning), is hiring a Senior Data Scientist to build predictive models and insights that drive real action for colleges and businesses using their Engagement Platform. In this role, you will own predictive modeling, unstructured data analysis, AI evaluation, and analytics surfaces that make data accessible and trustworthy. You'll work on a lean team alongside a Senior Data Engineer and Senior Analytics Engineer, with significant autonomy in deciding which signals to pursue and which requests to decline. Key responsibilities include: - Finding and shipping predictive signals: Investigate signals across student engagement, outcomes, and partner health. Build predictive models that reach partners through systems they already use, with output that drives real action. Set thresholds against actionable alert volumes and evaluate performance across student populations, not just in aggregate. - Analyzing conversational and unstructured data: Apply embeddings, clustering, and classification to conversational data, support records, and other unstructured sources to surface themes and emerging concerns. Build text analysis foundations that make unstructured data retrievable and useful to AI tooling. - Owning AI data tooling and evaluation: Manage in-warehouse AI configuration, verified queries, prompts, and agent tooling. Build and maintain the AI evaluation framework with precise rubrics, defensible sampling, and reporting that shows whether model or prompt changes actually improved outcomes. - Making work reusable: Equip internal teams with data and analysis for strategic partners, favoring generalizable work over one-offs. Document reasoning, assumptions, and tradeoffs in the internal knowledge base. Work in dbt/code alongside engineers. This is not a role focused on training models full-time. You'll need judgment about when not to build a model, when an analysis or conversation is better, and when to decline requests. You'll set the methodology bar for modeling and evaluation work and be the person others trust when they aren't sure whether to trust a number. Requirements: - 5+ years building predictive models that were actually used, including several years owning the problem end-to-end (not just being handed it), with demonstrated judgment around calibration, threshold-setting, and feature leakage detection. - Practical NLP experience: text classification, clustering, embeddings, or similar applied work (calling an LLM API alone is insufficient). - Strong SQL as a primary tool, not just a way to get data into a notebook. - Working Python for modeling and analysis. - Solid applied statistics with judgment about which methods fit the question and when data can't support a conclusion. - Modern cloud warehouse experience (Snowflake, BigQuery, Databricks, or similar), especially with in-warehouse AI or agent tooling. - Excellent written and verbal communication; ability to hand findings to non-technical colleagues and have them act on them. - Care about how predictions get used and willingness to check whether models work equally well across different populations (e.g., part-time vs. full-time students). - Comfort with ambiguity and honesty about uncertainty; preference for saying "the data can't answer this" over confident answers that fall apart later. - Comfortable working in version control with code review for reproducible analysis and models. - Experience productionizing model output into operational workflows. - Experience scheduling and monitoring recurring jobs in production with judgment to reach for maintainable tooling. - Track record of choosing what to work on: turning ambiguous business goals into scoped projects, making prioritization cases, and being accountable for impact. Nice to have: - Transformation tooling (dbt or similar), dimensional modeling, or analytics engineering exposure. - Familiarity with AI evals or prompt evaluation. - Modern BI tool experience (Sigma, Looker, Hex, Tableau, or similar). - Experience working closely with analytics or data engineers where your models depended on others' tables. - Linguistics or computational linguistics background. - EdTech, higher education, or student success background.

Similar roles