SlipstreamJobsFresh Startup & VC-Backed Jobs

AI Research Engineer - Datadog AI Research (DAIR)

Datadog - Paris, Île-de-France, France - In-office - posted 2025-08-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Join Datadog AI Research (DAIR) as an AI Research Engineer, partnering with Research Scientists to transform cutting-edge AI research into production systems. You'll build the data pipelines, tooling, and infrastructure that enable rapid iteration from prototype to deployment. Datadog AI Research focuses on two high-impact areas: (1) World Models for Observability—training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events to power forecasting, anomaly detection, root cause analysis, and autonomous agents; and (2) Trained Agents for Observability—post-training models to operate autonomously in SRE incident response, code repair, security response, and infrastructure optimization. Key responsibilities include: - Design and operate multimodal data pipelines, training and evaluation infrastructure, benchmarks, and internal tooling - Implement models, run experiments at scale, and profile for reliability, performance, and cost - Build simulation environments and replay infrastructure for agent training and evaluation - Orchestrate distributed training and distributed RL using Ray, including scheduling, scaling, and failure recovery - Establish rigorous automated benchmarks and regression tests for world model predictions and agent performance - Collaborate with Research Scientists, Product, and Engineering teams to integrate capabilities into Datadog products - Contribute to research publications at top-tier venues (NeurIPS, ICLR, ICML) and produce high-quality code and documentation You bring depth in distributed computing, RL infrastructure, and ML systems for training and inference at scale. You're proficient in Python, familiar with systems languages (Rust, C++, Go), and comfortable with modern cloud and data infrastructure. You have practical experience implementing and operating ML training systems (PyTorch, JAX), including containerization, orchestration, and GPU acceleration. Experience with large-scale model training frameworks (Megatron-LM, DeepSpeed, VeRL) and techniques like SFT, RLHF, and efficient inference is expected. You can explain technical trade-offs clearly and have supported or contributed to research publications. Bonus experience includes observability/SRE/security domains, bridging research prototypes to production with foundation models or RL agents, GPU programming and CUDA optimization, production data pipelines, and building simulation environments for agent training.

Similar roles