SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company has ~200 employees distributed worldwide with no physical office. You'll join the Data side of the AI team, responsible for building and operating the infrastructure that powers large-scale, high-quality dataset collection for model training.
In this role, you'll own multiple aspects of the data pipeline: sourcing new audio data, operating and extending cloud infrastructure (GCP, Terraform), and collaborating with research scientists to optimize the cost/throughput/quality tradeoff. You'll work on petabyte-scale data ingestion systems and help shape the AI team's dataset roadmap to support next-generation consumer and enterprise products.
Key responsibilities include:
- Identifying and integrating new audio data sources into the ingestion pipeline
- Operating and extending GCP-based cloud infrastructure managed with Terraform
- Working closely with scientists to deliver richer datasets at scale and lower cost
- Contributing to strategic decisions on the AI team's data roadmap
You should have a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, and strong proficiency with Python/bash scripting, Docker, and Infrastructure-as-Code. Experience with web crawlers and large-scale data processing is a plus. The role values scrappiness, adaptability, and strong communication in an asynchronous, distributed environment.