SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer, Data Platform

Matterworks - Somerville, MA, United States - Hybrid

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Matterworks is building foundation models that identify and interpret the "dark matter" of biochemical biology—the vast majority of molecular signals that mass spectrometers detect but remain unidentified. The company's Large Spectral Models apply AI to make metabolites, lipids, and peptides legible and predictable, transforming R&D across the life sciences. As Senior Software Engineer on the Data Platform team, you will design and build the connective tissue of Matterworks' data infrastructure. You'll own end-to-end data pipelines that serve both ML research and product, reporting to the Head of Engineering and collaborating daily with machine learning researchers, scientists, and product teams. Key responsibilities include: • Build and scale data contracts: design and implement pipelines that other teams depend on, handling multiple petabytes of data efficiently. • Serving and cost at scale: architect systems to acquire, store, and serve data affordably as it grows, optimizing data layout, featurization throughput, Kubernetes orchestration, and cost visibility. • Labels and enrichment: transform raw spectrometry data into usable datasets with consistent schemas, trustworthy metadata, and documented definitions. Scale scientific labels from studies to individual spectra and features. • Quality gates: automate validation checks that enable capability growth without regression, promoting only passing data. • Interfaces: own the surfaces (SDKs, tools, APIs) that AI, chemistry, product, and agent teams call to build and access datasets. • Operations and data rights: ensure reliable platform operations meeting SLAs while maintaining data provenance, security, and compliance with customer requirements. You bring significant production experience building data systems and pipelines. You're proficient in Python and SQL for large-scale processing, and experienced with Kubernetes-native orchestration (Argo Workflows, Metaflow, EKS) and modern data lake technologies (Apache Iceberg, Parquet, DuckDB, Terraform). Experience with Airflow or Dagster transfers well. You've designed stable identifiers for large, evolving datasets and built validation gates that prevent bad data from publishing. You have hands-on experience integrating LLMs or agents into production data pipelines, including evaluation loops, gold sets, and cost tracking. You use AI coding tools daily with healthy skepticism about their correctness on production data. You own work end-to-end—from design through running, validated systems—including unglamorous tasks like fixing malformed datasets and debugging failed publishes. Comfort with scientific data formats (mzML, RDKit, ProteoWizard) is valuable. Above all, you thrive in early-stage environments where autonomy, learning, and solving novel scientific challenges matter more than rigid process.

Similar roles