SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pencil is an AI-powered SaaS platform that uses generative AI to create advertising content at scale—video, adaptations, and multiple formats orchestrated by an agentic system called Scribble. The company's mission is to democratize AI-driven creative production for businesses of all sizes, not just enterprises.
This Senior Product Manager role owns the end-to-end evaluation and quality measurement function for Pencil's advertising creative output. You will be responsible for building reliable measurement systems in a subjective domain where there is no single "right answer." Your charter spans internal systems that measure whether agents, skills, and workflows produce client-ready work, as well as external evidence used in customer RFPs, creative enablement packages, and adevals.ai—Pencil's public evaluation site.
Key responsibilities include:
- Owning the Evals & Quality charter: defining the durable problem statement, success metrics, and explicit scope boundaries.
- Defining core quality metrics (first-pass-right, coverage, regression detection) and driving continuous improvement.
- Building the judging methodology: designing rubrics for creative quality and maintaining calibration between automated judges and human raters over time.
- Designing and scaling the human QC mechanism: determining how humans review output, how their judgments feed the automated layer, and how the system scales.
- Curating and versioning the golden dataset to remain representative across formats and client contexts.
- Creating self-service eval loops that other product teams can use independently.
- Owning adevals.ai as a product, ensuring it meets a high design bar for credibility.
- Providing the evidence layer for commercial work: RFP responses and creative enablement packages.
- Partnering with Agent Architects on client-specific quality bars through configuration and rubrics.
- Running the full product lifecycle with written hypotheses before every ship and post-launch audits 14 days after release.
You will thrive in this role if you move fast with imperfect information, stay deeply human-centered in your design, act like an owner, keep solutions simple, and delight users through thoughtful polish.
Success will be measured by: first-pass-right rate (the lead metric), eval coverage across skills and formats, judge-human calibration agreement, golden dataset health, commercial evidence (adevals.ai live and used in wins), and learning velocity (post-launch reviews against stated hypotheses).
REQUIREMENTS:
- Demonstrated ownership of a quality, evals, or ML measurement problem with its own roadmap and users.
- Shipped measurement in a subjective domain: experience defining quality and maintaining reliable signal at scale.
- Direct hands-on experience with golden datasets, rubric design, judge calibration, and inter-rater reliability—including failures and how you recovered.
- A product decision you derived from evidenced customer need in an unfamiliar domain, with a bias toward discovery over pre-formed solutions.
- Ownership of a metric that changed decisions outside your own team.
- Time spent in front of clients presenting or defending methodology to buyers.