SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 208,000 - 315,000 / annual
Snorkel AI is hiring a Senior/Staff Software Engineer to join the early ML & Research Engineering team. The company, which raised $350M Series E at $3.5B valuation in September 2026, is scaling to meet demand for its data-centric AI platform.
You will work on frontier AI data generation and evaluation, making the process faster, cheaper, and more rigorous. As an early team member, you'll study how frontier-grade data is generated and evaluated, form and validate hypotheses against production data, and ship solutions at scale. You'll shape the discipline's direction, standards, and the team that grows around it.
Key focus areas include:
- Efficient agentic evals: cutting costs through adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating
- AI model routing: routing eval and judge calls to the cheapest model that meets quality bars, with fallback and monitoring
- Fine-tuned small models: fine-tuning and serving open-weight models (LoRA and parameter-efficient methods) where they match frontier quality
- Predictive difficulty: building models to estimate task difficulty for frontier systems before rollouts
- Measurement for AI data: building golden datasets, quantifying LLM-as-judge accuracy and calibration, ensuring reproducible quality
- Research to production: turning research prototypes into reusable, configurable components for forward-deployed engineers and researchers
This is a pure ML/AI engineering role with no infrastructure plumbing—every problem is an open ML or LLM problem. Your work directly impacts the speed, quality, and cost of the data frontier AI is built on.
REQUIREMENTS:
- 5+ years building production ML or software systems with end-to-end ownership from prototype to production
- Hands-on experience running LLM or ML workloads in production, comfortable reasoning about non-deterministic systems
- Deep grounding in statistics and experimentation: experiment design, hypothesis testing, sampling, confidence intervals
- Strong Python and software engineering fundamentals (testing, code review, system design)
- Experience designing evaluations and interpreting results rigorously
- Habit of finding high-impact problems before assignment; clear communication with researchers, engineers, and business partners
NICE TO HAVE:
- Fine-tuning and serving open-weight models, judging when smaller models meet quality bars
- Building LLM evaluation or experimentation platforms, model gateways, or routing systems
- Experience with agentic workloads, benchmarks, or RL environments
- Record of taking research to production (publications, open-source work, shipped research-driven features)
- MS or PhD in Computer Science, Machine Learning, Statistics, or related field