SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 140,000 - 200,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees in a fully distributed setting, Speechify is building AI-powered audio solutions that support people with learning differences like dyslexia, ADD, low vision, and autism.
You'll join the Data side of the AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research. This is a hands-on engineering role where you'll own critical infrastructure and data pipelines.
Key responsibilities include: sourcing new audio data and integrating it into ingestion pipelines; operating and extending cloud infrastructure (GCP, Terraform) for data ingestion; collaborating with research scientists to optimize the cost/throughput/quality frontier; and helping shape the AI team's dataset roadmap for next-generation consumer and enterprise products.
You should have a BS/MS/PhD in Computer Science or related field, 5+ years of software development experience, strong proficiency with Python/bash in Linux, hands-on experience with Docker and Infrastructure-as-Code, and familiarity with major cloud providers (GCP preferred). Experience with web crawlers and large-scale data processing is a plus. The role demands adaptability, strong communication, and comfort operating in a fast-moving, scrappy environment.
Speechify offers competitive salaries ($140k–$200k base), equity, bonus, and a laid-back, asynchronous culture with a hands-off management approach. You'll work on a product that directly impacts millions and operates at the intersection of AI and audio.