SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, news articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees distributed globally and no physical offices, Speechify operates a fully remote, asynchronous culture.
This role is part of Speechify's AI team, specifically focused on the data infrastructure and acquisition side. You will own all aspects of data collection to support model training operations, building high-quality datasets at petabyte scale while optimizing for cost efficiency.
Key responsibilities include:
- Identifying and integrating new audio data sources into the ingestion pipeline
- Operating and extending cloud infrastructure (GCP, Terraform) for data ingestion
- Collaborating with research scientists to optimize the cost/throughput/quality frontier, delivering richer datasets at scale and lower cost
- Working with the AI team and leadership to define the dataset roadmap for next-generation consumer and enterprise products
You should have a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, and strong proficiency in Python/bash scripting, Docker, Infrastructure-as-Code, and cloud platforms (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. The ideal candidate is scrappy, adaptable, and communicates clearly in writing and verbally.
Speechify offers a fast-growing environment with entrepreneurial culture, hands-off management, competitive salaries, and the opportunity to impact millions of users, particularly those with learning differences like dyslexia, ADD, low vision, and autism.