SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Reflection AI is a research lab building open foundational models to make intelligence accessible to everyone. This role is a Member of Technical Staff focused on multilingual data infrastructure and quality.
You will design and operate large-scale multilingual data pipelines, handling sourcing, cleaning, deduplication, language identification, and script normalization across high- and low-resource languages. You'll define and enforce quality standards for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.
Key responsibilities include:
- Design and run scientific experiments to advance understanding of scaling large language models and multilingual data efficiency
- Lead small research projects independently while collaborating on larger initiatives
- Build evaluation sets and diagnostics that expose where model behavior degrades by language, register, or domain
- Work with pre-training, mid-training, and post-training teams to deliver measurable improvements in multilingual capability
- Navigate trade-offs between research objectives and practical engineering constraints
This is a high-ownership role in a small, fast-moving team where you'll help define the company's future and the frontier of open foundational models.
REQUIREMENTS:
- Strong software engineering fundamentals and comfort processing web-scale datasets in distributed environments
- Experience building large-scale data pipelines for language models, machine translation, speech, or search (ideally covering multiple languages)
- Fluency or working proficiency in at least one language other than English, with genuine curiosity about linguistic differences
- Rigorous, measurement-first mindset: you prove data changes helped rather than assuming
- Ability to navigate ambiguity and operate with high ownership
- Comfortable with research-to-engineering trade-offs
NICE TO HAVE:
- Graduate degree (MS or PhD) in Computer Science, Machine Learning, or related discipline
- Published work in multilingual NLP, low-resource languages, or evaluation methodology
- Experience with annotation vendor management or crowdsourced data operations at scale
- Familiarity with tokenizer design and effects on non-Latin scripts