SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 30,000 - 90,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company has been recognized by Google (Chrome Extension of the Year) and Apple (2025 Design Award for Inclusivity) and operates as a fully distributed team of ~200 people with employees from Amazon, Microsoft, Google, Stanford, and other top companies.
You will join the Data side of Speechify's AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte-scale and low cost through tight integration of infrastructure, engineering, and research.
Key responsibilities:
- Source and integrate new audio data sources into the ingestion pipeline
- Operate and extend cloud infrastructure for data ingestion (GCP, Terraform)
- Collaborate with AI scientists to optimize the cost/throughput/quality frontier, delivering richer datasets at scale and lower cost
- Work with the AI team and leadership to shape the dataset roadmap for next-generation consumer and enterprise products
Required qualifications:
- BS/MS/PhD in Computer Science or related field
- 5+ years of industry software development experience
- Proficiency with bash/Python scripting in Linux environments
- Professional experience with Docker and Infrastructure-as-Code (GCP preferred)
- Strong communication skills
Preferred: experience with web crawlers and large-scale data processing workflows.
The role offers competitive salary, equity, a hands-off management approach, and the opportunity to impact a product used by millions, including people with learning differences like dyslexia, ADD, low vision, and autism.