SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Engineer

Tessera Labs - San Jose, CA, USA - In-office - posted 2026-08-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Tessera Labs is building an AI-powered enterprise transformation platform that helps large companies modernize their systems, processes, and data in weeks rather than years. The company raised a $60M Series A led by Andreessen Horowitz. As a Research Engineer, you will own the post-training and inference machinery that turns hypotheses about agent behavior into production models. This role bridges research and systems engineering: you build the infrastructure (training stacks, evaluation harnesses, RL environments) while Research Scientists own the research agenda. Key responsibilities include: - Building and scaling the post-training stack: supervised fine-tuning, preference optimization, and reinforcement learning for long-horizon tool use and enterprise system transformations. RL is central to this role. - Designing memory and context machinery for long-horizon agents: what agents retain across multi-step runs, how it's structured, retrieved, and compacted. - Building the representation layer (ontologies, knowledge graphs) that agents reason over, derived from real enterprise systems. - Creating data generation and curation pipelines: synthetic landscapes, transformation traces, tool-call trajectories, and curriculum infrastructure for training on enterprise systems no public model has seen. - Implementing RL environments: sandboxed execution harnesses where changes can be applied, verified, and scored automatically. - Building and running offline evaluation infrastructure for long-horizon agentic behavior, including trajectory-level scoring and reproducible task suites. - Running end-to-end experiments: design, launch, debug, analyze, and distinguish signal from noise. - Optimizing training and inference throughput: kernels, parallelism, memory, batching, serving—with special attention to long-context handling. - Taking training results to production: quantization, serving configuration, rollback paths. - Establishing standards for reproducibility, experiment tracking, and result hygiene. The role is uniquely interesting because much of the task space is verifiable: transformations either produce systems that build, pass regression suites, and behave equivalently, or they don't. This creates real reward signals rather than preference models. You'll work with open-weight models on rented compute clusters, operating in a compute-constrained environment where spending is justified by results.

Similar roles