SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company has ~200 employees distributed worldwide with no physical office.
You'll join the Data side of Speechify's AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research.
Key responsibilities:
- Source new audio data and integrate it into the ingestion pipeline
- Operate and extend cloud infrastructure for data ingestion (GCP, Terraform)
- Collaborate with research scientists to optimize the cost/throughput/quality frontier
- Work with AI team and leadership to define the dataset roadmap for next-generation consumer and enterprise products
Required qualifications:
- BS/MS/PhD in Computer Science or related field
- 5+ years of industry software development experience
- Proficiency with bash/Python scripting in Linux
- Professional experience with Docker and Infrastructure-as-Code
- Experience with at least one major cloud provider (GCP preferred)
- Strong communication skills
Desirable:
- Web crawler experience
- Large-scale data processing workflow expertise
The company emphasizes a fast-growing, entrepreneurial environment with hands-off management, competitive salaries, and asynchronous culture. The role directly impacts accessibility for people with learning differences including dyslexia, ADD, low vision, and autism.