SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Engineer, LangSmith Engine

LangChain - San Francisco, CA, USA - Hybrid - posted 2026-08-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

LangChain is hiring a Research Engineer to join the LangSmith Engine team, which builds a proactive agent engineer that analyzes production traces, identifies failures, recommends fixes, and prevents issues from recurring. The Engine is a production system that helps developers improve AI agents at scale. In this role, you will study real agent failures and build benchmarks that capture what matters most. You'll design and run experiments to improve agent performance across models, prompting, context, tools, orchestration, and agent strategies. This includes exploring post-training and fine-tuning techniques when they can meaningfully improve capabilities, quality, or cost. You'll turn successful experiments into production improvements, working closely with engineers and researchers to measure impact and prevent regressions. You'll also help define the ML roadmap and technical direction for improving Engine agents and mentor other engineers through strong technical leadership. A key aspect of this role is understanding production engineering and system-level tradeoffs—improving an agent means not just maximizing benchmark performance but also understanding impact on cost, latency, reliability, and scalability. You bring 4+ years of ML/AI research experience (or closely related field) with a Master's or Ph.D. in a relevant scientific field. You have hands-on experience with LLMs and AI agents, including analyzing model behavior and improving real-world performance. You're skilled at designing benchmarks, evaluations, and experiments for AI/ML systems and know how to tell whether a change actually made an agent better. You have strong software engineering skills with a track record of taking ideas from research prototype to measurable production impact. You have maximum agency and strong research judgment—you can identify high-impact problems, work through ambiguity, move quickly, and communicate findings clearly. Nice-to-have qualifications include a Ph.D. in Machine Learning, Computer Science, or Physics; hands-on experience with LLM-as-a-judge, automated graders, synthetic data generation, or human evaluation; experience with reinforcement learning, preference optimization, SFT, RLHF/RLAIF, or other post-training techniques; experience optimizing LLM agents for cost, latency, or task efficiency in production; and experience with model serving, inference optimization, distributed systems, or GPU infrastructure. LangChain has raised $125M at Series B from top-tier investors and serves 6,000+ active LangSmith customers, including 5 of the Fortune 10 and 35% of the Fortune 500. The company offers competitive compensation including base salary, equity, and benefits such as medical, dental, vision coverage, flexible vacation, 401(k), and meals on in-office days.

About LangChain

AI / Data / Infrastructure — developer platform and framework for building LLM applications and agents.

Similar roles