SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Statsig team is seeking a Machine Learning Engineer to lead the technical direction for ML-powered experimentation and insights capabilities. The Statsig platform provides experimentation, feature rollout, dynamic configuration, and analytics systems that enable OpenAI's product, engineering, research, and go-to-market teams to ship with speed, safety, and evidence.
This is a 0-to-1 role where you will build production systems that learn from privacy-protected product and experimentation data to generate evidence-backed insights and support decision-making. You will work across ML modeling, retrieval and LLM systems, statistical methods, simulation, data and training pipelines, backend services, and user- and agent-facing product experiences. The core challenge is making each insight and prediction traceable, calibrated, useful, and safe enough to influence real product decisions.
Key responsibilities include:
- Setting and executing the technical roadmap for Generative Insights and Predictive Experimentation from prototypes through production adoption
- Building cross-experiment learning systems that retrieve and synthesize historical experiments, detect recurring effects, and generate hypotheses with clear evidence
- Developing predictive models and simulation workflows to estimate impact, affected segments, regression risk, and uncertainty before live experiments
- Creating high-quality datasets and feature/retrieval pipelines with strong lineage, freshness, privacy, and data-quality controls
- Establishing rigorous evaluation through offline benchmarks, backtests, calibration, drift monitoring, and prediction-to-outcome comparisons
- Turning models into durable product, API, and agent workflows that move from insight to experiment design to measured learning
- Partnering with data science and product teams on experiment design, causal inference, and sequential decision-making
- Building reliable services and intuitive workflows so ML capabilities are understandable to teams making high-stakes decisions
- Providing technical leadership across engineering, product, data science, and research partners
The team is based in Bellevue and works in-person to move quickly, solve ambiguous problems together, and stay close to product teams. You will support teams across ChatGPT, Codex, model measurement, consumer experiences, business subscriptions, developer products, and shared infrastructure.
REQUIREMENTS:
- Led ambiguous 0-to-1 production ML products where success was measured by better real-world decisions, not only offline model metrics
- Strong hands-on experience across the ML lifecycle: dataset design, training/adaptation, evaluation, deployment, monitoring, and iteration
- Depth in one or more of: LLM and retrieval systems, ranking/recommendation, forecasting/anomaly detection, causal ML/experiment analysis, or simulation
- Strong software engineering fundamentals and ability to build high-quality production systems in Python, working across data, backend, and platform boundaries
- Strong grounding in machine learning, statistics, computer science, or related field through formal study or equivalent practical experience
- Understanding of experimentation and statistical reasoning, especially the distinction between predictive accuracy and causal validity
- Treat calibration, uncertainty, provenance, privacy, and human review as product requirements
- Ability to translate ambiguous partner questions into product and technical roadmaps
- Experience building for internal power users and agents, making sophisticated ML capabilities clear and actionable
- Value in-person collaboration and interest in shaping a growing Bellevue-based team