SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior AI Product Manager, Code

Scale - New York, NY, USA - Hybrid - posted 2026-08-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Scale is seeking a Senior AI Product Manager to own and scale its Coding portfolio, which includes SWE-Bench Pro, SWE Atlas, and contributions to FrontierBench. This role sits at the intersection of product strategy, technical depth, and customer engagement, driving Scale's position as the leading AI data foundry for coding model training and evaluation. You will define the strategy and roadmap for Scale's coding data products, reinforcement learning environments, and agentic coding evaluations. Key responsibilities include owning the end-to-end product lifecycle from ideation through launch and scaling; extending benchmark franchises as models saturate current tasks; partnering with ML researchers and engineers to develop trustworthy task specifications and quality standards; driving infrastructure roadmaps for Harbor-native environments, execution sandboxes, and contributor tooling; establishing governance for data quality, contamination prevention, and reproducibility; and growing the expert contributor network of professional software engineers. You will work directly with frontier AI labs and enterprise customers to understand model failure modes, gather feedback, and influence investment decisions. You'll collaborate across AI Product Management, ML Research, Engineering, Operations, and Go-To-Market teams to convert benchmark credibility into durable, revenue-generating product lines. Additional responsibilities include identifying new coding capability areas (long-horizon agentic work, repo-scale refactoring, debugging, code review), managing external partnerships across the coding ecosystem, and tracking adoption and business outcomes to guide resource allocation. The ideal candidate combines strong product judgment, hands-on technical depth in software engineering, operational rigor, and customer-facing experience. You should be able to read code, reason about codebases, and hold your own with senior engineers and ML researchers. Familiarity with how coding models are trained and evaluated, including post-training methods, agentic scaffolds, and coding benchmarks, is expected.

Similar roles