SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 30,000 - 110,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company was named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees across a fully distributed, asynchronous organization, Speechify is at the intersection of AI and audio accessibility.
You'll join the Data side of the AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research.
Key responsibilities:
- Source new audio data and integrate it into the ingestion pipeline
- Operate and extend cloud infrastructure (GCP, Terraform) for data ingestion
- Collaborate with research scientists to optimize the cost/throughput/quality frontier, delivering richer datasets at scale and lower cost
- Work with the AI team and leadership to shape the dataset roadmap for next-generation consumer and enterprise products
You should have a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, and proficiency with Python/bash scripting, Docker, and Infrastructure-as-Code on a major cloud provider (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. The role requires strong communication skills and the ability to adapt to changing priorities in a fast-moving environment.