SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees distributed globally and no physical office, Speechify operates a fully remote, asynchronous culture.
This role joins Speechify's AI team on the Data Infrastructure & Acquisition side, responsible for building and operating the systems that collect, ingest, and manage high-quality datasets at petabyte scale to power next-generation text-to-speech and AI models.
Key responsibilities include:
- Sourcing and integrating new audio data sources into the ingestion pipeline
- Operating and extending cloud infrastructure (GCP, Terraform) for data pipelines
- Collaborating with research scientists to optimize the cost/throughput/quality frontier of datasets
- Working with the AI team and leadership to define the dataset roadmap for consumer and enterprise products
The ideal candidate has a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, and strong proficiency in Python/bash scripting, Docker, Infrastructure-as-Code, and cloud platforms (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. The role values scrappiness, adaptability, and strong communication in an entrepreneurial environment.
Speechify offers competitive salaries, a hands-off management approach, and the opportunity to impact a product used by millions, including people with learning differences like dyslexia, ADD, low vision, and autism.