SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 30,000 - 100,000 / annual
Speechify is a consumer-facing text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. Speechify operates as a fully distributed, 100% remote company with ~200 employees across engineering, AI research, and product teams.
This role is part of Speechify's AI team, specifically focused on the data infrastructure and acquisition side. You will own all aspects of data collection to support model training operations, building high-quality datasets at petabyte scale while optimizing for cost efficiency. The role combines infrastructure engineering with research collaboration to power next-generation text-to-speech models.
Key responsibilities include: sourcing new audio data and integrating it into ingestion pipelines; operating and extending cloud infrastructure (GCP, Terraform-managed) for data pipelines; collaborating with AI scientists to optimize the cost/throughput/quality tradeoff; and working with leadership to define the AI team's dataset roadmap for consumer and enterprise products.
You should have 5+ years of software development experience, strong proficiency in Python/bash scripting and Linux environments, hands-on experience with Docker and Infrastructure-as-Code, and professional experience with a major cloud provider (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. The ideal candidate holds a BS/MS/PhD in Computer Science or related field and can juggle multiple priorities while communicating effectively across technical and non-technical stakeholders.