SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Data Infrastructure

Cohere - San Francisco, CA, United States - Hybrid - posted 2026-09-11

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Cohere is a leading enterprise AI company building cutting-edge foundation models and end-to-end products for businesses. The Data Infrastructure team is responsible for the storage and data movement layer underlying every model training run, serving petabytes of training data and model checkpoints to keep thousands of GPUs busy across multiple training clusters. In this role, you will design, build, and operate the distributed storage system that feeds model training and evaluation. You'll run this system on Kubernetes clusters at petabyte scale, working closely with researchers and training infrastructure teams to translate their data access patterns into throughput, latency, and durability requirements. You'll tackle networking, I/O, and consistency challenges involved in moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as key success metrics. This is a foundational engineering opportunity to build a critical system from the ground up at a scale few teams have tackled. You'll be a key contributor solving infrastructure problems that directly enable Cohere's model training capabilities. Cohere is remote-friendly with offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul. The company offers comprehensive benefits including health/dental coverage, RRSP/401K matching, 100% parental leave top-up for up to 6 months, 6 weeks paid vacation, education stipends, and a $500 home office stipend. REQUIREMENTS: - Strong storage fundamentals, including replication, consistency, caching, and data lifecycle management - Strong coding ability in Python and/or Go (willingness to learn the other required) - Hands-on experience running stateful systems on Kubernetes, including Persistent Volumes, CSI drivers, and StatefulSets - Hands-on experience with cloud object storage (e.g., S3) and POSIX-style filesystems BONUS QUALIFICATIONS: - Experience with parallel or HPC filesystems such as Weka, VAST, or Lustre - Familiarity with data-loading and checkpointing patterns used in large-scale model training - Interest in understanding how large models are trained and evaluated, including trajectory and eval data patterns

About Cohere

AI / Data / Infrastructure — enterprise generative AI models and tooling for businesses.

Similar roles