SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ElevenLabs is seeking a Research Engineer to build and operate large-scale, distributed web crawling systems that source high-quality training data for frontier AI models. The role focuses on discovering, fetching, and extracting data across billions of web pages reliably and efficiently.
Key responsibilities include:
- Designing and scaling distributed web crawlers that handle billions of pages with high reliability and efficiency
- Solving complex crawling challenges: content extraction from messy HTML, deduplication at web scale, freshness/recrawl strategies, and politeness/rate-limit handling
- Building targeted crawling pipelines to identify and extract high-value data sources including audio, video, and multilingual content, converting them into clean training-ready datasets
- Creating infrastructure and tooling that enables researchers to request, monitor, and explore newly crawled web data quickly and reliably
Required experience:
- Hands-on background building and scaling web crawlers or scraping systems, ideally for machine learning training data pipelines
- Strong distributed systems engineering skills at scale (Kubernetes, queue-based architectures, or custom pipelines processing billions of documents)
- Ability to autonomously evaluate data quality, coverage, and compliance; building tooling to measure these metrics
ElevenLabs operates with a high-velocity, impact-driven culture emphasizing rapid experimentation, lean autonomous teams, and minimal bureaucracy. The company uses AI-first approaches across engineering and operations. Founded in January 2023, ElevenLabs has raised $781M and reached an $11B valuation, serving millions of users and thousands of businesses including Deutsche Telekom and Meta. The team includes IOI medalists and ex-founders. While remote globally, optional office locations are available in London, New York, San Francisco, and Warsaw.