SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sunset is a rapidly scaling business process automation company that partners with frontier AI labs to provide de-identified, proprietary enterprise data for model training. Founded to help startups wind down, the company has pivoted into a high-value data licensing business, growing from $0 to multi-eight-figure run rate in months with backing from top-tier investors including Floodgate and Afore.
As Sunset's first Data Scientist focused on evaluation, you will own the measurement and quality assurance infrastructure for the company's core product: transforming sensitive internal enterprise data into de-identified datasets that preserve structure and utility. This is a zero-to-one role at the intersection of data science, AI systems, and production operations.
Your core responsibility is establishing how the team knows whether data quality is actually improving. You will design and build evaluation datasets, experiments, quality metrics, and feedback loops that expose hidden failures, accelerate model and pipeline improvements, and give stakeholders confidence in deliverables. You'll work hands-on with Python and SQL to construct evaluation corpora, study failure patterns, design rigorous comparisons, calibrate human and model-based judgments, and translate findings into actionable decisions.
Key challenges include: determining whether improvements in aggregate metrics (e.g., NER scores) actually reduce sensitive information misses across diverse data modalities and providers; distinguishing genuine quality gains from label noise, sample bias, or shifted workloads; identifying hidden failures in specific customer segments, entity types, languages, or formats that aggregate results mask; and designing the smallest credible experiments to support ship/revise/stop decisions.
You will partner closely with Machine Learning Engineers (who own model behavior), Product Engineering, Data Engineering, Security, Quality teams, and domain experts. While ML owns changing models, you own the credibility of evidence used to validate whether changes actually made data safer or more useful. Success means the team has decision-grade baselines for priority quality claims, improvements are judged by consequential slices and failure costs, and evaluation becomes faster and more repeatable without sacrificing rigor or independence.