SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Scale AI is seeking a Senior Frontier Agents Engineer to bridge cutting-edge AI research and production deployment. You'll work directly with enterprise customers across finance, healthcare, manufacturing, media, and telecommunications to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software.
In this role, you'll own the full lifecycle of modern AI systems: designing reasoning and agent architectures, building retrieval and memory systems, developing predictive models that work alongside LLMs, running rigorous experiments and ablation studies, shipping production systems into enterprise environments, and measuring business impact through online experimentation.
You'll design and deploy production AI agents leveraging the latest advances in large language models, reasoning, retrieval, memory, and tool use. You'll architect intelligent systems that combine LLMs, traditional ML, structured knowledge, and deterministic software into reliable production workflows. You'll engineer customer intelligence layers, retrieval pipelines, and memory systems that allow agents to reason over large, heterogeneous enterprise data, and develop multi-agent systems that coordinate reasoning, planning, tool execution, and human oversight.
Experimentation and evaluation are core to the role. You'll own the full experimentation lifecycle from hypothesis generation to production rollout, design rigorous evaluation frameworks using offline benchmarks, online A/B experiments, golden datasets, and human evaluation, and continuously evaluate newly released frontier models to determine where they meaningfully improve quality, latency, reliability, or cost.
Production quality is paramount. You'll build AI systems with strong emphasis on reliability, observability, latency, safety, and cost. You'll design agent guardrails, fallback strategies, tracing, and monitoring pipelines that enable safe deployment in high-stakes environments, and build human-in-the-loop workflows that effectively combine AI automation with expert oversight.
You'll partner directly with enterprise customers to understand their business, data, and operational challenges, translate ambiguous problems into production AI architectures, rapidly prototype and validate ideas, and identify reusable patterns that become core capabilities across deployments.