SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
DeepL is seeking a Senior Research Scientist to join the Foundation Model Task Adaptation (FMTA) team in London. This role focuses on developing and deploying cutting-edge research in reinforcement learning and post-training for large language models at scale.
You will design, implement, and deploy state-of-the-art reinforcement learning pipelines, post-train large multi-modal models to align them with human intent, and enable general capabilities such as reasoning. Your responsibilities span the entire research lifecycle: from idea conception and theoretical modeling through prototyping, ablation studies, and production deployment. You'll collaborate deeply with Engineering, ML Platform, and HPC teams to deliver robust model updates to users, and build external collaborations with academic and industrial partners.
The ideal candidate has a strong technical background with deep practical experience in Python and modern ML frameworks (PyTorch, TensorFlow, or JAX). You should have expertise in deep reinforcement learning techniques (RLHF, RLAIF, RLVR) and hands-on experience scaling and deploying LLMs or foundation models in real-world systems. A master's degree, PhD, or equivalent industry experience in mathematics, physics, computer science, or related fields is required. You'll demonstrate a track record of leading self-directed research projects that deliver tangible results beyond academic exercises.
DeepL offers a hybrid work schedule (three days in office per week) with flexible hours, a diverse globally-distributed team of 1,000+ employees across 90+ nationalities, and monthly Hack Fridays for passion projects. The company, founded in 2017, serves over 200,000 business customers and millions of individuals across 228 markets with its Language AI platform.