SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - London

H Company - London, United Kingdom - Hybrid - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

H builds computer-use agents and the models behind them, deployed through a managed API and integrated into enterprise workflows by forward-deployed engineers. This role owns the evaluation framework that measures agent performance across benchmarks—currently ~50, scaling to 100–200 soon. The framework orchestrates, runs, and observes benchmarks spanning web apps, desktop applications, and command-line interfaces. Researchers use evaluation results to choose model checkpoints and decide on releases; product and forward-deployed engineers use them to measure agents on customer workflows. You will: - Support researchers and forward-deployed engineers integrating new benchmarks, streamlining the process each time. - Define and enforce standards for how benchmarks enter the framework. - Optimize scheduling and observability to prevent cluster idle time and queue buildup. - Ensure reproducible results across trials so release decisions rest on reliable numbers. - Work across whatever stack benchmarks require—cluster tuning, browser extensions, desktop environments. - Engage with customers (individual developers to large enterprises) to understand what they need measured, then automate measurement so results flow into harnesses and models. In the first 3 months, you'll help integrate 5 benchmarks and start addressing framework bottlenecks. By 6 months, you'll own one subsystem (scaling, observability, or a benchmark group) and a release will ship on your numbers. By 12 months, you'll be H's go-to expert on the entire evaluation system. You'll report to Ceiran Chapman, VP Engineering, and work closely with H's researchers and forward-deployed engineers. REQUIREMENTS: - 5+ years backend development with production Python as core; comfortable using coding agents to move faster without sacrificing quality. - Built test, QA, or evaluation tooling that other teams depended on; care deeply about correctness. - Operated distributed systems on Kubernetes in public cloud (AWS experience most valuable). - Built and shipped end-to-end systems including APIs (REST or GraphQL) and integrations with external services. - Knowledge of relational and non-relational databases, message queues (SQS, RabbitMQ, Kafka). - Instrument systems from the start with metrics, tracing, and monitoring. STRONGER CANDIDATES ALSO HAVE: - Experience measuring LLM quality or building agents. - Docker and VM workload packaging and deployment. - Familiarity with Temporal, Dask, FastAPI, PostgreSQL, Grafana, or Datadog. - Automated web or desktop software using Playwright, Selenium, or custom browser extensions. - Set standards for other engineers through code review, design review, or mentoring. No machine learning background required; H will support you in learning that domain. If you match most but not all criteria, apply anyway.

Similar roles