SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, news articles, and websites into audio. The company has been recognized by Google (Chrome Extension of the Year) and Apple (2025 Design Award for Inclusivity) and operates as a fully distributed, 100% remote organization with ~200 employees.
You'll join the Data side of Speechify's AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research. This is a hands-on role where you'll directly impact the next generation of Speechify's consumer and enterprise AI products.
Key responsibilities include:
- Sourcing new audio data and integrating it into the ingestion pipeline
- Operating and extending cloud infrastructure for data ingestion (GCP, Terraform)
- Collaborating with research scientists to optimize the cost/throughput/quality frontier
- Working with the AI team and leadership to shape the dataset roadmap
You'll need a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, and proficiency with Python/bash scripting, Docker, and Infrastructure-as-Code. Experience with web crawlers and large-scale data processing is a plus. The role requires strong communication skills and the ability to adapt to changing priorities in a fast-moving environment.
Speechify offers competitive salaries, a laid-back asynchronous culture, and the opportunity to work on products that directly impact people with learning differences including dyslexia, ADD, low vision, and autism.