SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
H builds computer-use agents and the models behind them, deployed through a managed API and integrated into enterprise workflows by forward-deployed engineers.
This role owns the evaluation framework that measures agent performance across benchmarks—currently ~50, scaling to 100–200 soon. The framework orchestrates, runs, and observes benchmarks spanning web apps, desktop applications, and command-line interfaces. Researchers use evaluation results to choose model checkpoints and decide on releases; product and forward-deployed engineers use them to measure agents on customer workflows.
You will:
- Support researchers and forward-deployed engineers integrating new benchmarks, streamlining the process each time.
- Define and enforce standards for how benchmarks enter the framework.
- Optimize scheduling and observability to prevent cluster idle time and queue buildup.
- Ensure reproducible results across trials so release decisions rest on reliable numbers.
- Work across whatever stack benchmarks require—cluster tuning, browser extensions, desktop environments.
- Engage with customers (individual developers to large enterprises) to understand what they need measured, then automate measurement so results flow into harnesses and models.
In the first 3 months, you'll help integrate 5 benchmarks and start addressing framework bottlenecks. By 6 months, you'll own one subsystem (scaling, observability, or a benchmark group) and a release will ship on your numbers. By 12 months, you'll be H's go-to expert on the entire evaluation system.
You'll report to Ceiran Chapman, VP Engineering, and work closely with H's researchers and forward-deployed engineers.
REQUIREMENTS:
- 5+ years backend development with production Python as core; comfortable using coding agents to move faster without sacrificing quality.
- Built test, QA, or evaluation tooling that other teams depended on; care deeply about correctness.
- Operated distributed systems on Kubernetes in public cloud (AWS experience most valuable).
- Built and shipped end-to-end systems including APIs (REST or GraphQL) and integrations with external services.
- Knowledge of relational and non-relational databases, message queues (SQS, RabbitMQ, Kafka).
- Instrument systems from the start with metrics, tracing, and monitoring.
STRONGER CANDIDATES ALSO HAVE:
- Experience measuring LLM quality or building agents.
- Docker and VM workload packaging and deployment.
- Familiarity with Temporal, Dask, FastAPI, PostgreSQL, Grafana, or Datadog.
- Automated web or desktop software using Playwright, Selenium, or custom browser extensions.
- Set standards for other engineers through code review, design review, or mentoring.
No machine learning background required; H will support you in learning that domain. If you match most but not all criteria, apply anyway.