SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Paris

H Company - Paris, France - Hybrid - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

H builds computer-use agents and the models behind them, deployed through a managed API and integrated into enterprise workflows by forward-deployed engineers. This role owns the evaluation framework that measures agent performance across benchmarks—currently ~50, scaling to 100–200 soon. The framework orchestrates, runs, and observes benchmarks spanning web apps, desktop applications, and command-line interfaces. Researchers use evaluation results to select model checkpoints and decide on releases; product and forward-deployed engineers use them to measure agents on customer workflows. You will: - Support researchers and forward-deployed engineers integrating new benchmarks, streamlining the process each time. - Define and enforce standards for how benchmarks enter the framework. - Optimize scheduling and observability to prevent cluster idle time while evaluation jobs queue. - Ensure reproducible results across trials so release decisions rest on reliable numbers. - Work across diverse stacks—cluster tuning, browser extensions, desktop environments—whatever a benchmark requires. - Engage with customers (individual developers to large enterprises) to understand what they need measured, then automate measurement so results flow back into harnesses and models. In the first 3 months, you'll help integrate 5 benchmarks and start addressing framework bottlenecks. By 6 months, you'll own one subsystem (e.g., scaling, observability, or a benchmark group) and ship a release on your numbers. By 12 months, you'll be the company expert on evaluation system design and trade-offs. You'll report to Ceiran Chapman, VP Engineering, and work directly with H's researchers and forward-deployed engineers. REQUIREMENTS: - 5+ years backend development with production Python as core; comfortable using coding agents to accelerate without sacrificing quality. - Built test, QA, or evaluation tooling that other teams depended on; care deeply about correctness. - Operated distributed systems on Kubernetes in public cloud (AWS experience most valuable). - Built and shipped end-to-end systems including APIs (REST/GraphQL) and integrations with external services. - Knowledge of relational and non-relational databases, message queues (SQS, RabbitMQ, Kafka). - Instrument systems from the start with metrics, tracing, and monitoring. STRONGER IF YOU HAVE: - Experience measuring LLM quality or building agents. - Packaged and run workloads in Docker and virtual machines. - Used Temporal, Dask, FastAPI, PostgreSQL, Grafana, or Datadog. - Automated web or desktop software with Playwright, Selenium, or custom browser extensions. - Set standards other engineers follow through code review, design review, or mentoring. No machine learning background required. If you match most but not all criteria, apply anyway.

Similar roles