SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 140,000 - 200,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees distributed globally and no physical office, Speechify combines engineering, AI research, and product development to serve users with learning differences and accessibility needs.
This role joins Speechify's AI team on the data infrastructure side, responsible for building and operating systems that collect, process, and manage high-quality datasets at petabyte scale to power next-generation text-to-speech models. The position bridges engineering and research, focusing on cost-efficient, large-scale data acquisition.
Key responsibilities include: sourcing new audio data streams and integrating them into ingestion pipelines; operating and extending cloud infrastructure (GCP, Terraform) for data collection; collaborating with research scientists to optimize the cost/throughput/quality tradeoff; and helping shape the AI team's dataset roadmap to support both consumer and enterprise product development.
Ideal candidates have 5+ years of software engineering experience, strong proficiency in Python/bash scripting and Linux environments, hands-on experience with Docker and Infrastructure-as-Code, and familiarity with at least one major cloud provider (GCP preferred). Experience with web crawlers and large-scale data processing workflows is valued. The role requires adaptability, strong communication, and comfort working asynchronously in a distributed team.
Speechify offers competitive compensation, equity participation, a hands-off management style, and the opportunity to impact a fast-growing intersection of AI and audio technology while supporting users with dyslexia, ADHD, low vision, and other learning differences.