SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff - Multilingual Data

Reflection AI - San Francisco, CA, USA - In-office - posted 2026-09-11

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Reflection AI is a research lab building open foundational models to make intelligence accessible to everyone. This role is a Member of Technical Staff focused on multilingual data infrastructure and quality. You will design and operate large-scale multilingual data pipelines, handling sourcing, cleaning, deduplication, language identification, and script normalization across high- and low-resource languages. You'll define and enforce quality standards for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks. Key responsibilities include: - Design and run scientific experiments to advance understanding of scaling large language models and multilingual data efficiency - Lead small research projects independently while collaborating on larger initiatives - Build evaluation sets and diagnostics that expose where model behavior degrades by language, register, or domain - Work with pre-training, mid-training, and post-training teams to deliver measurable improvements in multilingual capability - Navigate trade-offs between research objectives and practical engineering constraints This is a high-ownership role in a small, fast-moving team where you'll help define the company's future and the frontier of open foundational models. REQUIREMENTS: - Strong software engineering fundamentals and comfort processing web-scale datasets in distributed environments - Experience building large-scale data pipelines for language models, machine translation, speech, or search (ideally covering multiple languages) - Fluency or working proficiency in at least one language other than English, with genuine curiosity about linguistic differences - Rigorous, measurement-first mindset: you prove data changes helped rather than assuming - Ability to navigate ambiguity and operate with high ownership - Comfortable with research-to-engineering trade-offs NICE TO HAVE: - Graduate degree (MS or PhD) in Computer Science, Machine Learning, or related discipline - Published work in multilingual NLP, low-resource languages, or evaluation methodology - Experience with annotation vendor management or crowdsourced data operations at scale - Familiarity with tokenizer design and effects on non-Latin scripts

Similar roles