SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Wynd Network builds infrastructure that delivers massive amounts of web data to companies training the world's most powerful AI models. The company operates Grass, a bandwidth-sharing network powering a massive distributed crawler with unique access to high-quality public web data at global scale. On top of that, they've built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.
As a Machine Learning Engineer, you will join a small, innovative team and lead efforts to advance the company's capabilities, drive model development, and support the vision for a future where Grass is transformative in the internet's evolution. This is a fully remote role, though it requires sufficient overlap with EST business hours for effective team collaboration.
Key responsibilities include:
- Developing data processing pipelines and machine learning solutions for large-scale NLP and LLM applications, including improving the quality, filtering, and preparation of training datasets
- Designing and implementing pipelines for processing and analyzing large datasets
- Analyzing and interpreting complex time series data to provide actionable insights and solutions
- Designing, implementing, and maintaining data-driven models and algorithms
- Developing techniques for dataset curation to improve the quality and efficiency of AI training data
- Building scalable pipelines for filtering, deduplicating, and improving large-scale text datasets used for LLM training
- Collaborating with cross-functional teams to understand data needs and deliver timely solutions
- Ensuring data quality and integrity throughout all processes
- Utilizing Optical Character Recognition (OCR) technology to convert different types of documents into editable and searchable data
- Continuously researching and implementing best practices in data science and machine learning
- Contributing to the development and improvement of internal data processing tools and infrastructure
The company is lean, technical, and moves fast with no red tape or slow decision-making—just builders pushing to expand what's possible for open web data and AI. They prioritize low ego and high output, seeking people who make everyone around them better.
REQUIREMENTS:
- Bachelor's, Master's, or Doctoral degree in Data Science, Computer Science, Statistics, or a related field
- Minimum 3 years of work or research experience dealing with large datasets
- Strong coding skills in Python or other object-oriented programming languages
- Graduate-level knowledge of statistics, including hypothesis testing, regression analysis, and probability
- Excellent work ethic and ability to thrive in a fast-paced startup environment
- Strong problem-solving skills and attention to detail
- Good communication skills with ability to articulate complex data concepts to non-technical stakeholders
- Experience working in a high-output team
PREFERRED:
- Experience working with large-scale text datasets, NLP pipelines, or data preparation for LLM training
- Experience with text deduplication, dataset filtering, corpus curation, or data distillation