SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
DeepL is seeking a Research Manager to lead the Production Inference team, responsible for the systems that serve DeepL's language AI models reliably and efficiently at scale. This role combines people leadership with technical direction, overseeing research scientists and ML engineers working on performance-critical model serving systems.
Key responsibilities include leading and developing a high-performing team with strong development plans and feedback culture; owning the team's research and development roadmap for production inference systems while balancing near-term reliability with longer-horizon research bets; serving as the primary technical interface between Production Inference and adjacent functions (foundational models research, voice research, applied research, infrastructure, and product); driving reliability, efficiency, and cost performance of DeepL's model serving stack including strategic decisions around serving infrastructure evolution; operating with high autonomy in an environment with ambiguous or evolving requirements; and actively identifying, assessing, and recruiting research and engineering talent.
The Production Inference team sits at the intersection of research-grade technical ambition and production-grade operational discipline. The team owns the full model serving stack—from GPU-resident inference runtimes and deployment infrastructure through to the developer platform enabling other research teams to bring models to production. The team operates within the Research organisation with close ties to infrastructure and platform functions, and architectural decisions affect every product DeepL ships. The broader Research organisation publishes regularly at ACL, NeurIPS, and EMNLP.
Required qualifications include a PhD in Computer Science, Mathematics, Physics, or comparable quantitative discipline, or strong ML/systems background with equivalent research depth. Candidates should have strong foundation in production ML systems, inference optimisation, or model serving at scale, with direct experience in LLM inference, speculative decoding, quantisation, or serving infrastructure being a meaningful differentiator. Proven experience leading teams of researchers or ML engineers with track record of developing talent, maintaining delivery rigour, and balancing research quality with production reliability is essential. Comfort operating across the full model lifecycle from training through production deployment, monitoring, and efficiency improvement is required. Excellent communication skills and ability to translate complex technical direction into clear goals for technical and non-technical stakeholders are critical.