SlipstreamJobsFresh Startup & VC-Backed Jobs

QA Engineer-AI Native Quality

Newton Research - Boston, MA, United States - In-office - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 115,000 - 130,000 / annual

Newton Research is a fast-growing AI-native software startup building the next generation of the closed-loop media lifecycle. The company develops AI agents leveraging large language models and generative AI with specialized knowledge to generate actionable business insights for media planning, buying, and measurement. This QA Engineer role focuses on quality assurance for AI-native products where traditional testing approaches fall short. The core mission is to build and maintain evaluation frameworks for AI agent behavior, since agent outputs vary run-to-run and require probabilistic rather than deterministic assertions. Key responsibilities include: • Own the evaluation suite for Newton's agents: curate datasets of inputs and expected behaviors, build LLM-as-judge scoring systems, track regressions across model/prompt/skill/agent releases, and validate judges against human-labeled examples. • Test skill and prompt changes before shipping: ensure every edit to a skill, system prompt, or tool definition runs against relevant evals in CI with before/after comparisons visible to reviewers; gate releases on results. • Cover agent flows end-to-end: validate tool-call correctness, task completion, multi-turn coherence, blueprint creation, code generation, and scheduled tasks; flag wrong, empty, or silently degraded output. • Handle non-determinism rigorously: use repeated runs, pass-rate thresholds, and statistical analysis so flaky agent behavior is measured, not anecdotal. • Convert production and customer signals into permanent evals: mine logs, error tracking, and customer reports to ensure escaped bad-behavior cases become regression tests. • Probe AI-specific risks: test for prompt injection, data leakage across users/projects/permissions, hallucinated analytics output, and cost/latency regressions. • Automate repetitive checks: any manual regression test done twice becomes a Playwright test, eval, or agent workflow; drive down manual testing time each sprint. • Design agentic QA workflows in CI: build agents that run test suites, triage failures, draft defects, and re-verify fixes, with guardrails and cost limits you define. • Maintain exploratory testing: conduct hands-on testing across roles, feature flags, environments, SSO/connector authorization, and cross-feature scenarios. • Write machine- and human-readable bug reports; partner with engineering on testability (observability, seedable data, stable interfaces). You will work alongside an existing QA lead and own the evaluation and skill-change testing layer—the gap that currently has no safety net. You are a release-readiness partner making the judgment call "are we good?" but are not hired to line-by-line test Claude-driven PRs. REQUIREMENTS: • 4+ years in QA or test engineering on complex web products, ideally B2B SaaS shipping frequently • Hands-on experience building evals for LLM or agent products (datasets, rubrics, LLM-as-judge, regression tracking), or demonstrated ability to learn fast • Working knowledge of how agents work: prompting, skills, tool calling, context, RAG, orchestration—enough to pinpoint failure origins • Automation-first mindset; can demonstrate what you removed from manual processes • Python (or similar) for eval tooling and data checks; Playwright or similar for end-to-end tests • Statistical literacy: pass rates, variance, sample size for non-deterministic outputs • Daily user of AI coding and testing assistants with judgment to verify output • Strong exploratory testing instincts and excellent written communication • Nice to have: adtech, martech, or marketing analytics domain; SSO/SAML/OAuth flows; data-connector testing; eval or observability tooling experience

Similar roles