SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 30,000 - 160,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, documents, and web content into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees distributed globally and no physical offices, Speechify operates a fully remote, asynchronous culture.
This role joins the Data side of Speechify's AI team, responsible for all aspects of data collection and pipeline infrastructure to support model training operations. The team builds high-quality datasets at petabyte scale while optimizing for cost and quality.
Key responsibilities include:
- Sourcing new audio data sources and integrating them into the ingestion pipeline
- Operating and extending cloud infrastructure (GCP, Terraform) for data ingestion
- Collaborating with research scientists to optimize the cost/throughput/quality frontier of datasets
- Working with the AI team and leadership to define the dataset roadmap for next-generation consumer and enterprise products
Required qualifications:
- BS/MS/PhD in Computer Science or related field
- 5+ years of industry software development experience
- Proficiency with bash/Python scripting in Linux environments
- Professional experience with Docker and Infrastructure-as-Code on a major cloud provider (GCP preferred)
- Strong communication skills
Preferred experience includes web crawlers and large-scale data processing workflows. The role offers competitive salary, equity, and the opportunity to impact a fast-growing AI/audio intersection while supporting accessibility for users with learning differences.