SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company has ~200 employees distributed worldwide with no physical office. You'll join the Data side of the AI team, responsible for building and operating the infrastructure that powers large-scale, high-quality dataset collection for model training.
In this role, you'll own multiple aspects of data infrastructure: sourcing new audio data streams, operating and extending the cloud ingestion pipeline (GCP + Terraform), and collaborating with research scientists to optimize the cost/throughput/quality tradeoff. You'll work across the AI team and leadership to shape the dataset roadmap that powers next-generation consumer and enterprise products.
Key responsibilities include:
- Identifying and integrating new audio data sources into the ingestion pipeline
- Operating and extending cloud infrastructure on GCP using Terraform
- Working with scientists to deliver richer datasets at scale and lower cost
- Contributing to AI team strategy and product roadmap
You should have a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, strong proficiency in Python/bash scripting and Linux, hands-on experience with Docker and Infrastructure-as-Code, and familiarity with at least one major cloud provider. Experience with web crawlers and large-scale data processing is a plus. The ideal candidate is adaptable, communicates clearly, and thrives in a fast-moving, entrepreneurial environment.
Speechify offers competitive salaries, a hands-off management approach, and the opportunity to impact millions of users, particularly those with learning differences like dyslexia, ADD, low vision, and autism.