SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a text-to-speech platform used by over 50 million people to convert PDFs, books, Google Docs, articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees globally in a fully distributed setting, Speechify operates at the intersection of AI and audio accessibility.
You'll join the Data side of the AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte-scale and low cost through tight integration of infrastructure, engineering, and research.
Key responsibilities include: sourcing new audio data and integrating it into the ingestion pipeline; operating and extending cloud infrastructure for data ingestion (GCP, Terraform); collaborating with research scientists to optimize the cost/throughput/quality frontier; and working with AI leadership to shape the dataset roadmap for next-generation consumer and enterprise products.
You should have a BS/MS/PhD in Computer Science or related field, 5+ years of industry software development experience, proficiency with bash/Python in Linux environments, hands-on experience with Docker and Infrastructure-as-Code, and professional experience with a major cloud provider (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. Strong communication skills and ability to handle multiple priorities are essential.
Speechify offers a fast-growing, entrepreneurial environment with hands-off management, competitive salaries, and the opportunity to impact millions of users with learning differences including dyslexia, ADD, low vision, and autism.