SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Exa is an applied AI lab building a next-generation search engine powered by massive-scale infrastructure, state-of-the-art embedding models, and high-performance vector databases. The company currently powers search for Cursor, Cognition, HubSpot, and over 400,000 developers, with $350M in funding from Lightspeed, Benchmark, and a16z.
As a Research Engineer focused on Content Understanding, you will work on the foundational problem of search quality: understanding what web pages actually contain. Before any page can be retrieved, the system must parse it into content vs. structural elements, classify page type and topic, assess usability and trustworthiness, extract publication dates, evaluate quality, and detect semantic duplicates. This work spans classical document understanding and emerging challenges like credibility assessment, AI-generated content detection, and crawler-optimized pages.
You'll have significant autonomy to shape projects based on your interests and strengths. Example work includes: improving parsing robustness with measurable validation; building models to judge page quality with consensus-based supervision; modeling credibility and misinformation as classification problems; deduplicating web-scale content while preserving user-relevant distinctions; designing scalable classification and extraction pipelines; creating supervision strategies for unlabeled problems; and debugging search failures to their root cause in page-level predictions.
The role emphasizes that most wins come from thoughtful data and supervision design rather than architectural novelty, and that defining ground truth for novel problems is part of the job.
REQUIREMENTS:
- Graduate-level ML experience: Master's degree or PhD with at least 2 years of relevant experience, or exceptionally strong undergraduate background
- Ability to build transformers from scratch in PyTorch
- Experience training models with production cost constraints (models that must run efficiently at scale)
- Comfort building and analyzing large-scale datasets
- Ability to work on problems where ground truth is undefined and must be established as part of the solution
- Genuine interest in the problem of high-quality knowledge discovery and its importance for society