SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Scale AI is seeking a Staff Frontier Agents Engineer to bridge cutting-edge AI research and production deployment. You'll work directly with enterprise customers across finance, healthcare, manufacturing, media, and telecommunications to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software.
In this role, you'll own the full lifecycle of modern AI systems: designing reasoning and agent architectures, building retrieval and memory systems, developing predictive models that work alongside LLMs, running rigorous experiments and ablation studies, shipping production systems into enterprise environments, and measuring business impact through online experimentation.
Key responsibilities include architecting intelligent systems that combine LLMs, traditional ML, structured knowledge, and deterministic software into reliable production workflows; designing and deploying production AI agents leveraging latest advances in large language models, reasoning, retrieval, memory, and tool use; developing multi-agent systems that coordinate reasoning, planning, tool execution, and human oversight; and translating frontier AI research into production systems by rapidly evaluating new models, prompting techniques, and agent architectures.
You'll design rigorous evaluation frameworks using offline benchmarks, online A/B experiments, golden datasets, regression suites, LLM-as-a-Judge, and human evaluation. You'll run controlled experiments to understand the contribution of different models, prompts, retrieval strategies, and reasoning techniques. You'll continuously evaluate newly released frontier models and determine where they meaningfully improve quality, latency, reliability, or cost.
Production quality is paramount: you'll build systems with strong emphasis on reliability, observability, latency, safety, and cost. You'll design agent guardrails, fallback strategies, tracing, monitoring, and evaluation pipelines that enable safe deployment in high-stakes environments. You'll partner directly with enterprise customers to understand their business challenges, translate ambiguous problems into production AI architectures, and rapidly prototype and validate new ideas.