SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
CuspAI is an AI-driven materials discovery company using machine learning to accelerate the discovery of breakthrough materials for energy, clean water, computing, and carbon capture. The company's founding team includes world-leading researchers in AI, chemistry, and engineering.
As Materials Knowledge Architect, you will lead the design of CuspAI's scientific data infrastructure, owning how heterogeneous external data is represented, modeled, and ingested into the platform. You will work at the intersection of materials science and data engineering, ensuring that data from diverse sources—industrial partners, research institutions, instrument vendors, and commercial providers—is transformed into high-quality, ML-ready assets.
Key responsibilities include:
**Data Architecture & Schema Design**: Design data models representing chemical, structural, and materials property information. Establish frameworks for data quality, validation, deduplication, lineage, and provenance. Leverage deep materials science expertise to maximize data quality and consistency from external sources.
**Ingestion & Pipeline Execution**: Lead ingestion of data from partners in varied formats and standards. Build robust pipelines that normalize and harmonize data. Identify, evaluate, and assess new data sources, partnerships, and data generation opportunities including computational (DFT, molecular dynamics) and experimental campaigns.
**Partner & Interdisciplinary Collaboration**: Work directly with external data providers to understand data formats, define schema mappings, and resolve quality and interoperability issues. Partner with research and engineering teams to translate modeling needs into ingestion requirements. Communicate opportunities to improve data models and processes.
Required qualifications: PhD in Chemistry, Physics, Materials Science, or related discipline (or equivalent research experience). Proven experience designing data models and schemas for complex scientific domains. Demonstrable experience ingesting and harmonizing data from multiple external sources with differing formats and quality levels. Expert knowledge of materials and chemical data sources (ICSD, Cambridge Structural Database, NOMAD, Materials Project). Working knowledge of Python and SQL with experience building data pipelines. Familiarity with materials science toolkits (pymatgen, ASE, Pandas/Polars). Firsthand understanding of experimental data production—lab workflows, instrument outputs, metadata challenges. Strong communication and collaboration skills.
Bonus: Industry experience in materials, chemicals, energy, or deep-tech. Familiarity with ETL/ELT pipelines, schema validation, workflow orchestration (Airflow, Dagster), and cloud infrastructure.