SlipstreamJobsFresh Startup & VC-Backed Jobs

Foundation Model Data Engineer

Sciforium - San Francisco, CA, United States - In-office - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary high-efficiency serving platform. The company is backed by multi-million-dollar funding and direct sponsorship from AMD, with hands-on support from AMD engineers. You will lead the strategy, creation, and curation of massive datasets powering foundation models. This role owns the end-to-end data lifecycle—from raw web-scale crawling to fine-grained human-alignment datasets that define model behavior. You will design taxonomies, filtering heuristics, and post-training pipelines ensuring models excel in reasoning, safety, and multimodal understanding. Key responsibilities include: - Foundation Dataset Strategy: Own end-to-end creation of pre-training datasets for LLMs, defining the optimal mix of web data, code, books, and technical papers to optimize downstream model performance. - Petabyte-Scale Curation: Design and implement sophisticated pipelines for data cleaning, exact/fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data. - Post-Training & Alignment Data: Lead development of high-quality post-training datasets, including Supervised Fine-Tuning (SFT) instructions, multi-turn dialogues, and preference modeling data (RLHF/DPO). - Multimodal Expansion: Drive acquisition and processing of vision and video data, navigating complexities of multimodal alignment, video compression, and temporal data consistency. - High-Performance Engineering: Develop high-throughput data processing scripts using Python, leveraging multiprocessing and multithreading to handle massive-scale ingestion and transformation. - Data Profiling & Analysis: Conduct deep-dive statistical analysis on training corpora to identify biases, gaps in knowledge, and quality regressions, ensuring mathematically balanced model "diet". - Synthetic Data Generation: Design pipelines to generate high-reasoning synthetic data augmenting gaps in natural datasets, utilizing existing models for data labeling and refinement. REQUIREMENTS: Must-haves: - 5+ years of industry experience in Data Science or Machine Learning with proven track record building and managing datasets for foundation models - Deep proficiency in Python with expert-level skills in high-performance code, multiprocessing, multithreading, and efficient memory management for large-scale data tasks - Petabyte-scale experience: demonstrated work with petabyte-scale datasets directly used to train production-grade LLMs or Large Vision Models - Dataset reconstruction experience: building massive LLM training sets from scratch, including raw web crawls (e.g., Common Crawl) and specialized domain data - Post-training expertise: hands-on experience building datasets for RLHF, DPO, and multi-turn instruction following, including management of human-labeling workflows and quality gold-sets - Data tooling mastery: frameworks such as Spark, Ray, or high-performance data-loading formats (e.g., WebDataset, Parquet) Nice-to-haves: - Computer Vision (CV) curation experience building large-scale image or video datasets from scratch (e.g., LAION-style pipelines) - Multimodal crawling familiarity with large-scale crawling of multimodal data and associated challenges of video processing, codecs, and compression - Taxonomy design experience in designing complex labeling schemas for reasoning, coding, and mathematical benchmarks - Research background: Master's or PhD in a quantitative field with focus on data-centric AI or information retrieval

Similar roles