SlipstreamJobsFresh Startup & VC-Backed Jobs

QA Engineer (AI Systems)

Nexxa.ai - Toronto, ON, Canada - In-office - posted 2026-08-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Nexxa is building autonomous AI systems for heavy industries—manufacturing, large-scale infrastructure, and logistics. The company enables machines and operations to think, decide, and act independently in complex, legacy environments. This role owns quality assurance for Nexxa's AI agent systems—products that plan, call tools, and execute multi-step actions autonomously in high-stakes industrial settings. This is not traditional UI testing. You will design evaluation frameworks for non-deterministic, tool-using systems; build golden datasets; catch regressions in reasoning quality; and stress-test agent behavior under adversarial and real-world edge-case conditions. You will work closely with ML engineers, backend engineers, and Forward Deployed Engineers to define what "good" looks like for an agent operating in industrial environments, then build infrastructure and processes to measure it continuously. Key responsibilities include: - Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence. - Develop golden datasets and labeled test sets, including edge cases, ambiguous inputs, and adversarial prompts specific to industrial contexts. - Define and track quality metrics beyond accuracy—groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations. - Build automated pipelines that run evals on every model, prompt, or tool-integration change, integrated into CI/CD. - Conduct structured red-teaming and adversarial testing (prompt injection, jailbreaks, tool misuse, unsafe actions) with security teams. - Test agent behavior across the full action loop—planning, tool selection, tool execution, error recovery, and final output. - Investigate and triage failures where root cause could be the model, prompt, tool/API, or orchestration logic. - Partner with ML and backend engineers to translate eval failures into actionable, reproducible bug reports. - Establish quality bars and sign-off criteria for new agent capabilities before customer deployment. - Mentor other engineers on testing strategies for probabilistic, LLM-driven systems. - Advocate for testability and observability in agent architecture from day one. REQUIREMENTS: - 5+ years in QA/SDET roles with demonstrated ownership of test strategy for complex systems. - Hands-on experience testing LLM-based products, chatbots, or AI agents; understanding of why traditional deterministic test assertions break down for generative systems. - Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or track record of building your own. - Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling. - Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks. - Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift. - Familiarity with LLM-specific failure modes: hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism. - Comfortable operating in ambiguity—defining what "correct" means when there is no single right answer. - Strong written communication skills for turning fuzzy quality signals into clear, actionable findings. PREFERRED: - Experience with human-in-the-loop evaluation workflows (labeling pipelines, inter-rater reliability, rubric design). - Background in ML/data science sufficient to read model evals and statistical significance. - Experience red-teaming or doing adversarial/security testing on ML systems. - Familiarity with observability/tracing tools for LLM applications (LangSmith, Arize, Langfuse, Weights & Biases). - Experience testing AI systems in industrial, IoT, or operational technology (OT) environments. - Prior experience setting up eval infrastructure from scratch at a startup or fast-moving team.

Similar roles