SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Krea is building next-generation AI creative tools focused on making AI intuitive and controllable for creatives. The company has raised over $83M from top-tier investors including a16z and Bain Capital, and recently launched Krea 2, their first foundation model built from scratch for aesthetic diversity and stylistic control.
As an ML Researcher, you will work on large-scale image and video diffusion model training experiments. Your responsibilities include training diffusion models on large GPU clusters, fully optimizing and profiling distributed training runs across model architectures, kernels, data loading, memory constraints, and communication patterns. You'll implement and improve distributed training strategies including FSDP, CP, SP, TP, and EP, continuously improving model quality through data curation, architecture design, training pipeline optimization, experiment structuring, and evaluation design.
You'll debug distributed training errors and implement fault tolerance solutions, identifying hardware issues (bad GPUs, NVLink, Infiniband components) and monitoring numerical errors and NCCL issues. You'll ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance.
The ideal candidate has a proven track record working with image or video models at scale (publications or open-source contributions preferred), strong PyTorch proficiency with deep understanding of its internals, and solid background in distributed training paradigms. You should be comfortable profiling and debugging large distributed training runs, analyzing traces to identify bottlenecks, and have good knowledge of low-precision training in FP8, NVFP4, and MXFP8. A solid understanding of the full diffusion model training pipeline—pretraining, midtraining, preference optimization, and reinforcement learning—is essential.
You'll work in a goal-oriented research environment where you're comfortable with underspecified goals, breaking ambiguous research objectives into concrete requirements and execution plans. The role requires good research taste, bias toward simplicity and methods that scale well, rapid iteration capability, and willingness to get hands-on with data pipeline design. The team includes musicians, designers, visual artists, and engineers, and the company emphasizes fast iteration, execution speed, and demonstrated interest in the creative space. Work is full-time and in-person at the North Beach office in San Francisco.