SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior/Staff FDE - CUA

Snorkel AI - New York, NY, United States - Hybrid - posted 2026-08-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Snorkel AI is hiring a Forward Deployed Engineer focused on Computer Use Agents (CUA) to lead technical execution of complex customer engagements involving AI agents that operate computers, browsers, and software environments to complete multi-step tasks. You will translate ambiguous product and model challenges into robust task environments, datasets, evaluators, and delivery plans that improve agent reliability and performance. Working across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery—you will identify patterns across engagements and turn successful approaches into reusable capabilities and product improvements. Key responsibilities include: **Computer Use Agents, Data, and Evaluation:** Design and build task environments, datasets, and evaluation workflows for agents operating across browsers, desktop applications, terminals, and other software interfaces. Translate customer goals and agent failure modes into representative multi-step tasks with clear success criteria. Develop data-generation, validation, and QA pipelines for multimodal and agentic training/evaluation data. Build automated evaluators and measurement frameworks to assess task completion, correctness, robustness, and efficiency. Diagnose agent failures across planning, tool use, perception, and state management; turn findings into improved tasks and data. Design and run experiments measuring how data, task design, and evaluation changes affect agent performance. Deliver production-grade task suites, datasets, and evaluation assets. **Forward Deployed Engineering & Customer Partnership:** Lead technical workstreams from solution design through production delivery, navigating ambiguity and making sound technical decisions. Build and iterate on solutions addressing customer needs. Rapidly prototype and productionize solutions across models, agent frameworks, APIs, and custom applications. Communicate technical tradeoffs and experimental results clearly to stakeholders. Serve as a trusted technical partner resolving complex blockers. **Technical Leadership & Scale:** Identify recurring patterns across engagements and turn successful solutions into reusable frameworks and best practices. Define and improve technical standards for agent task design, environment reliability, and evaluation. Partner with engineering, research, and product teams to influence platform capabilities. Lead technical design reviews and provide guidance to other engineers. Stay current with emerging agentic-AI and evaluation techniques. Required qualifications: 5+ years in machine learning engineering, software engineering, applied AI, forward deployed engineering, or similar technical role. Strong Python skills and experience building reliable production software, data, or ML systems. Hands-on experience building, evaluating, or deploying LLM-based or agentic systems, including computer-use agents. Strong understanding of experimentation and evaluation, including LLM-as-a-judge evaluation, metric definition, and using empirical results to guide decisions. Experience designing task environments, datasets, and verifiers for agents, including reward and verifier design. Experience building or working with APIs, automation, and web applications.

Similar roles