SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 140,000 - 200,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company was named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees in a fully distributed setting, Speechify combines talent from major tech companies (Amazon, Microsoft, Google), leading researchers, and startup veterans.
You'll join the Data side of the AI team, responsible for all aspects of data collection supporting model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research.
Key responsibilities:
- Source and integrate new audio data streams into the ingestion pipeline
- Operate and extend cloud infrastructure (GCP, Terraform) for data ingestion
- Partner with research scientists to optimize the cost/throughput/quality frontier, delivering richer datasets at scale and lower cost
- Collaborate with the AI team and leadership to define the dataset roadmap for next-generation consumer and enterprise products
Ideal candidates have a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, proficiency in bash/Python scripting on Linux, hands-on experience with Docker and Infrastructure-as-Code, and familiarity with major cloud providers (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. You should be comfortable juggling multiple priorities, adapting to change, and communicating clearly.
The company offers a fast-growing environment, entrepreneurial culture, hands-off management, competitive salary, equity, and the opportunity to impact millions of users with accessibility-focused technology.