SlipstreamJobsFresh Startup & VC-Backed Jobs

RE/RS, Data Understanding - Foundations

OpenAI - Zurich, Switzerland - In-office - posted 2026-08-19

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

The Data Understanding team at OpenAI is responsible for creating high-quality datasets and their quantized representations for use in large-scale model training. This includes synthesizing data, building vector quantization (VQ) representations, and handling processing, filtering, deduplication, quality control, and tokenization. In this role, you will advance how OpenAI builds and understands pretraining data at scale. You'll treat data quality and curation as core research problems, developing new methods to select, combine, and transform data. Your work will focus on creating datasets that improve model capabilities and designing rigorous experiments to understand how data choices and interventions affect model learning and downstream behavior. You'll work closely with frontier models and web-scale data to build evidence for which approaches work and why, then translate successful research into scalable data processing pipelines. This is a research-driven position where you'll own and drive your own research agenda, from problem selection through long-running work to measurable impact. Key expectations include a strong track record of new or improved ML ideas demonstrated through publications, projects, or applied research. You should be comfortable owning research agendas end-to-end and be excited by OpenAI's empirical, collaborative approach to research. Nice-to-have qualifications include thoughtfulness about AI's broader impact (privacy, provenance, data quality) and experience building high-performance deep learning or large-scale data processing systems. This role offers the opportunity to work on foundational problems in data understanding that directly impact the capabilities of cutting-edge AI systems.