SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cartesia is building foundational AI models using State Space Models (SSMs), a novel architecture invented by the founding team at Stanford AI Lab. The company is backed by leading VCs including Index Ventures, Lightspeed, and others.
As a Research Engineer in Data Infrastructure, you will own the systems and pipelines that power pretraining at Cartesia. Data is central to model quality, and this role bridges research and infrastructure to build the datasets and systems that enable cutting-edge language models.
Key responsibilities:
- Build and operate performant, scalable data processing infrastructure for acquiring, ingesting, and combining massive text datasets
- Design and operate scalable, high-throughput, reproducible data pipelines covering ingestion, preprocessing, filtering, deduplication, and augmentation
- Design and run ablation experiments to understand how data sources, processing choices, and mixture weights affect model quality
- Partner closely with research and infrastructure teams to co-design data loading, versioning, and experimentation pipelines
- Establish and enforce rigorous standards for data quality with tight feedback loops between dataset characteristics and model behavior
- Identify and source novel datasets; manage relationships and budgets with external data vendors and partners
Cartesia operates in-person offices in San Francisco, London, and Bangalore. The company culture emphasizes shipping fast, supporting each other, and maintaining high bars for quality and design.
Requirements:
- Hands-on experience with ML data infrastructure: training data pipelines, dataset versioning, large-scale data loading, and the interplay between data systems and model training/inference
- Strong modern engineering execution: clean, well-tested code, fluency with current tools, and judgment to pick the right tool for the problem
- Familiarity with building and evaluating datasets for generative models and working knowledge of how they're trained and used for inference
Nice-to-haves:
- Experience with large-scale data processing using parallel infrastructure such as Ray, Spark, or Kubernetes
- Experience with pretraining language models