SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anysphere, maker of Cursor (an AI-native code editor), is hiring Software Engineers to build the data systems behind frontier coding models' pretraining. The role spans three specialized teams:
Data Quality Team: Own the pipeline from raw internet-scale data to training-ready tokens. Build high-throughput, fully telemetered data pipelines with end-to-end traceability. Train and ship models that classify, rank, filter, and clean data at extreme throughput while maintaining accuracy and speed. Design and run scaling-ladder experiments on data mixtures, repeatability, and quality depth. Partner with data acquisition and training teams to close the loop on what moves loss and downstream evaluations.
Data Platform Team: Build infrastructure and pipelines transforming raw data dumps into training-ready datasets. Improve speed, reliability, and developer experience of initial training data pipelines. Enable researchers to quickly experiment with new data sources, quality filters, taxonomies, multimodal data, and data mixes. Create clear signals for data quality, lineage, freshness, and pipeline health.
Crawling Team: Build and scale web crawling systems that discover, schedule, fetch, and parse high-quality documents across the open web. Improve URL seeding, scoring, and fair host scheduling. Raise crawl success and parsing quality by defeating antibot failures and improving extractors. Debug and harden complex crawl infrastructure for availability, recovery, and ingestion lag. Work independently and alongside AI agents.
You'll treat data quality as both a systems and research problem—writing performance-critical code one week and designing careful experiments the next. The organization is flat and talent-dense, emphasizing truth-seeking, passion, creativity, and shipping code.
About Anysphere
AI / Data / Infrastructure — maker of Cursor, an AI-native code editor for software developers.