SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 160,000 - 210,000 / annual
Office Hours is an on-demand expert network connecting organizations with trusted experts across knowledge domains. The company is hyper-growth and profitable, with headquarters in San Francisco and offices in Brooklyn and Bangalore. They're backed by top marketplace investors from companies like DoorDash, Airbnb, and Affirm.
You'll build and run the platform behind AI model evaluations, working closely with the research team. Your responsibilities include preparing and maintaining benchmark datasets—owning data cleaning, preparation, format conversion, and ongoing maintenance while validating task completeness and consistency. You'll build and maintain evaluation pipelines that run consistently across model APIs and terminal agents, ensuring reproducible and comparable results.
You'll create lightweight, containerized environments and viewers for tasking and evaluating model performance on tool use. You'll support model experiments by helping fine-tune small open-source LLMs and comparing baseline against post-training performance. You'll develop the scoreboard and leaderboard—published views of results at model, benchmark, task, domain, and rubric levels.
Additionally, you'll build analysis tools to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements over time. You'll collaborate closely with researchers and engineers to ensure evaluation data and outputs are accurate and well-integrated into published work. You'll also build tooling for data creation and review, supporting expert annotation and data-generation projects through lightweight HTML viewers and internal tools.
You bring 4+ years of professional experience building and maintaining complex systems with strong Python skills. You write robust, maintainable code and dive deep into existing codebases. You have experience preparing, cleaning, and maintaining datasets with attention to data rigor. You're comfortable with Docker and building reproducible execution environments. You work well alongside researchers and scientists, translating methodology into working systems. Hands-on experience with AI evaluations or frameworks like Harbor, Terminal-Bench, or Inspect is a strong plus.