SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 30,000 - 120,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, news articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees across a fully distributed, 100% remote organization, Speechify is at the intersection of AI and audio accessibility.
You'll join the Data side of Speechify's AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte-scale and low cost through tight integration of infrastructure, engineering, and research.
Key responsibilities include:
- Sourcing new audio data and integrating it into the ingestion pipeline
- Operating and extending cloud infrastructure for the ingestion pipeline on GCP, managed with Terraform
- Collaborating with research scientists to optimize the cost/throughput/quality frontier, delivering richer datasets at scale and lower cost
- Working with the AI team and leadership to define the dataset roadmap for next-generation consumer and enterprise products
You should have a BS/MS/PhD in Computer Science or related field with 5+ years of industry software development experience. Required skills include bash/Python scripting in Linux, Docker, Infrastructure-as-Code, and hands-on experience with a major cloud provider (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. The ideal candidate is scrappy, adaptable to changing priorities, and communicates clearly both written and verbally.
Speechify offers a fast-growing, entrepreneurial environment with hands-off management, competitive salaries, stock options, and the opportunity to impact millions of users with accessibility-focused products.