SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cognichip builds AI-native tools for semiconductor engineering, combining large proprietary models, agentic workflows, and domain-specific intelligence to accelerate chip design, verification, and optimization. As Senior Data Scientist for AI Training Data, you will own the datasets that power the company's models—a critical leverage point for model quality.
Semiconductor engineering data is specialized, scarce, and often tightly licensed, fundamentally different from general text and code used to train most large language models. Your mission is to design curation, synthetic data generation, and quality-modeling workflows that transform raw technical material into production-ready training and evaluation datasets at scale.
Key responsibilities include: curating licensed and open-source technical datasets through collection, cleaning, annotation, and integration across the engineering lifecycle; building automated pipelines for sourcing, license classification, and normalization of public data; designing synthetic and augmented data generation workflows to meet model training demand; developing quality-modeling approaches including deduplication, contamination/leakage detection, license and PII screening, and difficulty/diversity scoring; writing large-scale distributed data processing jobs in partnership with a dedicated platform team; translating model weaknesses into targeted, well-sourced datasets; building and maintaining retrieval/embedding datasets for product features; running exploratory analysis to guide modeling, product, and go-to-market decisions; and establishing data governance practices around license provenance, documentation, and compliance.
You'll work closely with domain engineers to understand model needs, collaborate with the AI team to connect dataset improvements to measurable performance gains, and partner across engineering, product, and business teams to convert ambiguous requirements into concrete datasets.
Required: MS or PhD in Computer Science, Data Science, Statistics, or related field; 5–10 years of hands-on data science or ML data work with ownership of production datasets; expert Python and strong SQL with distributed computing experience (Spark); practical experience preparing text or code corpora for LLM training, fine-tuning, or evaluation; solid applied statistics and ML foundations with major framework experience (PyTorch, TensorFlow, scikit-learn); familiarity with modern data orchestration and versioned storage; working knowledge of data governance (licensing, provenance, confidential data handling); and strong ability to work with domain experts.
Preferred: exposure to hardware/engineering domain data; experience building retrieval systems (chunking, embeddings, vector databases); synthetic data generation using LLMs and agentic pipelines; annotation tooling and workflows; open-source contributions or published research in data-centric AI; or prior startup experience defining a data function.