SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior/Staff FDE - Synthetic Data Generation

Snorkel AI - New York, NY, United States - Hybrid - posted 2026-08-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Snorkel AI is seeking a Senior Forward Deployed Engineer to lead technical execution on synthetic data generation initiatives for enterprise customers and AI labs. You will translate complex model and data challenges into effective data strategies, designing and building scalable synthetic data generation, transformation, filtering, and evaluation pipelines. Your work spans the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. Key responsibilities include: designing LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets; building automated evaluators and quality frameworks to assess correctness, relevance, diversity, and coverage; designing experiments to measure synthetic data impact on model performance; and packaging production-grade datasets with standardized formats and documentation. You will lead technical workstreams from initial design through delivery, rapidly prototyping and productionizing solutions across models, data pipelines, APIs, and custom applications. Beyond individual engagements, you will identify recurring patterns across customer work and turn successful solutions into reusable pipelines, evaluators, and best practices. You'll partner with internal DaaS Engineering and Product teams to influence platform capabilities, lead technical design reviews, and mentor other engineers. This role requires staying current with emerging synthetic data, LLM evaluation, and data curation techniques. Required qualifications: 5+ years in machine learning engineering, data science, applied AI, or forward deployed engineering. Strong Python proficiency and production ML/data systems experience (Docker, cloud deployment on AWS/GCP/Azure). Hands-on experience with LLMs, building model-based applications and GenAI workflows, and integrating systems through APIs. Deep understanding of ML experimentation, evaluation metrics, and empirical decision-making. Demonstrated experience building synthetic data, data augmentation, or model-generated datasets. Familiarity with LLM evaluation techniques (LLM-as-a-judge, model-based evaluation, rubric-based evaluation). Ability to navigate ambiguous technical problems from definition through delivery with strong communication skills and customer collaboration experience.

Similar roles