SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Datadog's AI Platform organization is seeking a Senior Applied Scientist to lead the applied science direction for GenSim, the team building synthetic environments where Datadog's AI agents learn and train. GenSim creates fully instrumented, realistic applications that simulate production systems, inject controlled failures with known ground truth, and generate post-training data for Datadog's SRE and monitoring agents.
You will own the applied science methodology for this nascent effort, setting technical direction where none currently exists. Key responsibilities include: defining and measuring post-training data quality across correctness, representativeness, and difficulty; closing the realism gap between current simulated environments and messy real-world production systems; building scalable, production-grade synthetic environment systems (not research scripts); determining optimal application of this data in LLM post-training and agent evaluation; and collaborating cross-functionally with adjacent teams including Bits AI SRE, model training, and the evaluation pillar.
The open research questions are substantial: How do you ensure generated environments reflect the incomplete, imperfect telemetry real customers run? How do you make injected problems genuinely difficult and measure that difficulty? How do you control post-training data quality when correctness, representativeness, and difficulty create competing pressures? How do you evaluate non-deterministic agent trajectories end-to-end? Simulated agent environments for monitoring and SRE remain unsolved in open-source and published research; Datadog is the leading company in this space.
You bring 6+ years of applied science or ML engineering experience with demonstrated ability to set technical direction. You have hands-on expertise in LLM and agent post-training data creation, management, and quality control—this is the most critical requirement. You possess real domain expertise in LLMs and agentic systems (not classical ML fine-tuning), have evaluated agents or LLM applications, and can define success metrics before measuring. You are a strong production software engineer comfortable with Python, distributed systems, and shipping scalable infrastructure. You thrive in ambiguity, collaborate well across engineering and science teams, and are comfortable being the domain expert who decides what comes next. A PhD, MS, or equivalent research experience in a scientific field with strong applied mathematics grounding is required.