SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff AI/Machine Learning Engineer

Tonic.ai - Remote - Remote - posted 2026-07-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Tonic.ai builds data infrastructure for modern AI, generating synthetic environments for agent training and evaluation, and de-identifying enterprise data for safe use in model development. The company works with frontier AI labs and hundreds of enterprises including Fidelity, Comcast, eBay, and Vanguard. As a Staff AI/Machine Learning Engineer, you will design and build systems that generate longitudinally coherent synthetic environments for agent training and evaluation, including persona modeling, task generators, and verifiable ground truth. You'll build and maintain synthesis models that generate realistic replacement values at scale while preserving format, statistical distribution, and semantic consistency so de-identified data remains useful downstream. Key responsibilities include training and improving NER models for entity detection across free text, structured fields, and mixed enterprise data; building evaluation infrastructure that grades agent outcomes and discriminates between frontier models on real tasks; and fine-tuning open-weight models on Tonic-generated data to inform product and research direction. You will expand coverage into new domains, languages, and entity types, handling the long tail of formats and edge cases from real customer data. You'll own model evaluation across precision/recall for detection, utility preservation for synthesis, and outcome-level grading for agents. You'll optimize inference for efficient processing of large volumes of sensitive data within customer environments and partner directly with frontier labs and enterprise ML teams to turn hard data problems into shipped improvements. In this role, you'll set technical direction for a small senior team, raising the bar on rigor, reproducibility, and shipping. The work spans real range: one week might include building evaluation that separates best models from the rest, training synthesis models where both fidelity and downstream utility must hold, and improving entity detection on messy production data.

Similar roles