SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
G2 is the world's largest software marketplace, recently merged with Capterra, SoftwareAdvice, and GetApp to create the largest source of online data and software insights for B2B software buying decisions. The company serves 200M+ annual visitors and hosts 6M verified reviews.
As Software Engineering Director for Agentic Evaluations, you will lead the design and delivery of credible evaluations for AI agents offered by software vendors. This is a dual-focus role: 80% hands-on technical work building and enhancing the agent evaluation system, and 20% people leadership.
Key responsibilities include:
**Technical (80%):**
- Scope feasibility and effort for evaluating vendor-offered agents, identifying practical paths to credible evaluations
- Own the full stack of the core evaluation system, from admin and API surfaces to eval workflow dispatch
- Design evaluation system primitives that generalize across software verticals to enable reuse and faster onboarding
- Distill repeatable processes into AI skills or agents to increase evaluation velocity
- Monitor emerging frameworks and techniques for agent evaluation and industry-accepted distribution methods
- Collaborate with data science peers to develop rich, proprietary benchmarks
**Leadership (20%):**
- Facilitate growth and mentorship of a core engineering team
- Provide technical leadership and subject matter expertise in evaluations
- Socialize the use of evals across agent-oriented products in the wider organization
The role emphasizes "agent-first" ways of working, applying agentic engineering techniques to deliver production systems.
**Requirements:**
- 10+ years of professional programming experience in backend or full-stack environments
- 2+ years of direct experience managing engineers
- Expert-level proficiency in backend development using Python, Java/Kotlin, TypeScript/JavaScript, or Go; strong proficiency in associated frameworks like FastAPI or Node.js
- Direct experience creating evaluations against customer-facing agents, using agent trajectory trace data and rubrics to measure task completion, accuracy, correctness, or policy adherence
- Direct experience using frontier models from OpenAI, Anthropic, or Google in LLM-as-a-judge applications
- Regular use of coding agent harnesses (Claude Code, Codex, Opencode, Pi) as part of daily development workflow
- BS/BA degree in a related field
**Nice-to-have:**
- Experience with agent tool use via direct integration or MCP servers
- Knowledge of STATE-Bench, tau2-bench, or similar benchmark frameworks
- Experience with Playwright, browser-use, Chrome DevTools MCP for driving browser user flows
- Experience with durable execution frameworks (Temporal, DBOS, Cloudflare/Vercel Workflows) or agent sandboxing (AWS E2B, Daytona, Cloudflare/Vercel Containers)