SlipstreamJobsFresh Startup & VC-Backed Jobs

Pre-training Research Engineer

Sciforium - San Francisco, CA, USA - In-office - posted 2026-08-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD, the company is scaling rapidly to build the full stack powering frontier AI models and real-time applications. As a Pre-training Research Engineer, you will focus on model implementation, pre-training, and scaling while improving the quality of byte-native and multimodal foundation models. You'll build and iterate quickly on research ideas, contribute production-grade training code and infrastructure, and help deliver high-quality base models for real-world use cases at scale. Key responsibilities include training large byte-native and multimodal foundation models across massive, heterogeneous corpora; implementing and evaluating new model architectures, training objectives, and optimization methods; developing stable pre-training recipes and running scaling experiments for novel architectures; conducting ablations and analyzing training dynamics, model behavior, and base-model quality; and collaborating with data and distributed training engineers to improve training efficiency, reliability, and scalability. Required qualifications: 5+ years of experience in machine learning research or engineering with a proven track record of developing and pre-training large language or multimodal foundation models; strong general software engineering skills with the ability to write robust and performant training code; solid understanding of deep learning fundamentals and modern pre-training methods; ability to quickly implement research ideas and evaluate them using clear baselines, ablations, metrics, and analysis; hands-on experience running training workloads in GPU-based environments with familiarity in distributed training; and MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or related field. Nice-to-have qualifications include PhD in a related field, extensive experience with JAX, Flax, and XLA stack, experience with multi-node pre-training using FSDP, ZeRO, or Megatron, experience developing training recipes and scaling experiments, and experience owning end-to-end training and evaluation pipelines with monitoring and reproducibility.

Similar roles