SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary high-efficiency serving platform. The company is backed by multi-million-dollar funding and direct sponsorship from AMD, with hands-on support from AMD engineers.
You will lead the strategy, creation, and curation of massive datasets powering foundation models. This role owns the end-to-end data lifecycle—from raw web-scale crawling to fine-grained human-alignment datasets that define model behavior. You will design taxonomies, filtering heuristics, and post-training pipelines ensuring models excel in reasoning, safety, and multimodal understanding.
Key responsibilities include:
- Foundation Dataset Strategy: Own end-to-end creation of pre-training datasets for LLMs, defining the optimal mix of web data, code, books, and technical papers to optimize downstream model performance.
- Petabyte-Scale Curation: Design and implement sophisticated pipelines for data cleaning, exact/fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
- Post-Training & Alignment Data: Lead development of high-quality post-training datasets, including Supervised Fine-Tuning (SFT) instructions, multi-turn dialogues, and preference modeling data (RLHF/DPO).
- Multimodal Expansion: Drive acquisition and processing of vision and video data, navigating complexities of multimodal alignment, video compression, and temporal data consistency.
- High-Performance Engineering: Develop high-throughput data processing scripts using Python, leveraging multiprocessing and multithreading to handle massive-scale ingestion and transformation.
- Data Profiling & Analysis: Conduct deep-dive statistical analysis on training corpora to identify biases, gaps in knowledge, and quality regressions, ensuring mathematically balanced model "diet".
- Synthetic Data Generation: Design pipelines to generate high-reasoning synthetic data augmenting gaps in natural datasets, utilizing existing models for data labeling and refinement.
REQUIREMENTS:
Must-haves:
- 5+ years of industry experience in Data Science or Machine Learning with proven track record building and managing datasets for foundation models
- Deep proficiency in Python with expert-level skills in high-performance code, multiprocessing, multithreading, and efficient memory management for large-scale data tasks
- Petabyte-scale experience: demonstrated work with petabyte-scale datasets directly used to train production-grade LLMs or Large Vision Models
- Dataset reconstruction experience: building massive LLM training sets from scratch, including raw web crawls (e.g., Common Crawl) and specialized domain data
- Post-training expertise: hands-on experience building datasets for RLHF, DPO, and multi-turn instruction following, including management of human-labeling workflows and quality gold-sets
- Data tooling mastery: frameworks such as Spark, Ray, or high-performance data-loading formats (e.g., WebDataset, Parquet)
Nice-to-haves:
- Computer Vision (CV) curation experience building large-scale image or video datasets from scratch (e.g., LAION-style pipelines)
- Multimodal crawling familiarity with large-scale crawling of multimodal data and associated challenges of video processing, codecs, and compression
- Taxonomy design experience in designing complex labeling schemas for reasoning, coding, and mathematical benchmarks
- Research background: Master's or PhD in a quantitative field with focus on data-centric AI or information retrieval