SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Tessera Labs is building an AI-powered enterprise transformation platform that helps large companies modernize their systems, processes, and data in weeks rather than years. The company raised a $60M Series A led by Andreessen Horowitz.
As a Research Engineer, you will own the post-training and inference machinery that turns hypotheses about agent behavior into production models. This role bridges research and systems engineering: you build the infrastructure (training stacks, evaluation harnesses, RL environments) while Research Scientists own the research agenda.
Key responsibilities include:
- Building and scaling the post-training stack: supervised fine-tuning, preference optimization, and reinforcement learning for long-horizon tool use and enterprise system transformations. RL is central to this role.
- Designing memory and context machinery for long-horizon agents: what agents retain across multi-step runs, how it's structured, retrieved, and compacted.
- Building the representation layer (ontologies, knowledge graphs) that agents reason over, derived from real enterprise systems.
- Creating data generation and curation pipelines: synthetic landscapes, transformation traces, tool-call trajectories, and curriculum infrastructure for training on enterprise systems no public model has seen.
- Implementing RL environments: sandboxed execution harnesses where changes can be applied, verified, and scored automatically.
- Building and running offline evaluation infrastructure for long-horizon agentic behavior, including trajectory-level scoring and reproducible task suites.
- Running end-to-end experiments: design, launch, debug, analyze, and distinguish signal from noise.
- Optimizing training and inference throughput: kernels, parallelism, memory, batching, serving—with special attention to long-context handling.
- Taking training results to production: quantization, serving configuration, rollback paths.
- Establishing standards for reproducibility, experiment tracking, and result hygiene.
The role is uniquely interesting because much of the task space is verifiable: transformations either produce systems that build, pass regression suites, and behave equivalently, or they don't. This creates real reward signals rather than preference models. You'll work with open-weight models on rented compute clusters, operating in a compute-constrained environment where spending is justified by results.