SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff - ML Evals

Plastic Labs - New York, NY, USA - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 220,000 - 300,000 / annual

Honcho is a rapidly growing AI company building foundational technology for personal identity and agent alignment in the agentic world. The company has achieved 100x developer growth in six months, serves 50,000 prosumers and developers, and has processed half a trillion tokens. The vision extends beyond memory to creating 1:1 individual alignment for every human extending their cognition with AI—modeling any entity (person, agent, brand, team, NPC) to predict what that entity will do, believe, and choose, all without access to ground truth. In this role, you will own evaluation of Honcho end-to-end and the feedback loop it drives. Responsibilities include: **Design & Strategy**: Define evaluation scores and what "better" means for representations of entities that change over time. Keep definitions current as product and methods evolve. **Pipeline & Infrastructure**: Build the machinery (data ingestion, labeling, versioning, reruns, judges) that enables the team to ask new questions and get answers within a week. Own the infrastructure underneath, including building agents that run simulations at scale to increase iteration speed. **Execution & Analysis**: Run evaluations, harvest insights from traces and results, identify what's actually broken (not just what's easy to measure), and propose system changes that produce higher-fidelity representations. **Iteration & Discovery**: Find gaps in current evals, scope new experiments, run them, and document findings. You own the entire loop—question, pipeline, rerun, writeup—with nothing scoped off or waiting on others' sprints. Honcho is open source and mission-driven. The team values skepticism about measurements, rigorous methodology, and contribution to the broader open-source AI community. **Requirements**: - Trackable research experience (first or co-first author at NeurIPS, ICML, ICLR, or equivalent venues) OR independent open-source work of comparable quality - 3+ years in LLM evaluation, agentic optimization, or adjacent field - Strong Python experience leveraging agents for development - Familiarity with common data pipeline tooling (HuggingFace, Hydra, Kafka, SQL, etc.) - Ability to build production evals end-to-end and used by teams larger than yourself - Demonstrated skepticism about measurement validity; ability to identify when results are real vs. noise - Ability to construct measurements from nothing (no benchmark, labels, or ground truth) and argue for validity - Strong coding discipline: tests, issues, release cadence - Ability to grok complex systems quickly and build accurate mental models in hours - Belief in open-source methodology and transparency **Preferred**: Work on memory, identity, personalization, or agent evaluation; experience scoring representations that update over time; interests in cognitive sciences (linguistics, neuroscience, philosophy, psychology); active engagement with open-source AI community.

Similar roles