SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
H builds computer-use agents and the models behind them, deployed through a managed API and integrated into enterprise workflows by forward-deployed engineers.
This role owns the evaluation framework that measures agent performance across benchmarks—currently ~50, scaling to 100–200 soon. The framework orchestrates, runs, and observes benchmarks spanning web apps, desktop applications, and command-line interfaces. Researchers use evaluation results to select model checkpoints and decide on releases; product and forward-deployed engineers use them to measure agents on customer workflows.
You will:
- Support researchers and forward-deployed engineers integrating new benchmarks, streamlining the process each time.
- Define and enforce standards for how benchmarks enter the framework.
- Optimize scheduling and observability to prevent cluster idle time while evaluation jobs queue.
- Ensure reproducible results across trials so release decisions rest on reliable numbers.
- Work across diverse stacks—cluster tuning, browser extensions, desktop environments—whatever a benchmark requires.
- Engage with customers (individual developers to large enterprises) to understand what they need measured, then automate measurement so results flow back into harnesses and models.
In the first 3 months, you'll help integrate 5 benchmarks and start addressing framework bottlenecks. By 6 months, you'll own one subsystem (e.g., scaling, observability, or a benchmark group) and ship a release on your numbers. By 12 months, you'll be the company expert on evaluation system design and trade-offs.
You'll report to Ceiran Chapman, VP Engineering, and work directly with H's researchers and forward-deployed engineers.
REQUIREMENTS:
- 5+ years backend development with production Python as core; comfortable using coding agents to accelerate without sacrificing quality.
- Built test, QA, or evaluation tooling that other teams depended on; care deeply about correctness.
- Operated distributed systems on Kubernetes in public cloud (AWS experience most valuable).
- Built and shipped end-to-end systems including APIs (REST/GraphQL) and integrations with external services.
- Knowledge of relational and non-relational databases, message queues (SQS, RabbitMQ, Kafka).
- Instrument systems from the start with metrics, tracing, and monitoring.
STRONGER IF YOU HAVE:
- Experience measuring LLM quality or building agents.
- Packaged and run workloads in Docker and virtual machines.
- Used Temporal, Dask, FastAPI, PostgreSQL, Grafana, or Datadog.
- Automated web or desktop software with Playwright, Selenium, or custom browser extensions.
- Set standards other engineers follow through code review, design review, or mentoring.
No machine learning background required. If you match most but not all criteria, apply anyway.