SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 140,000 - 200,000 / annual
Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company has been recognized by Google (Chrome Extension of the Year) and Apple (2025 Design Award for Inclusivity) and operates as a fully distributed, 100% remote organization with ~200 employees worldwide.
You'll join the Data side of Speechify's AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research. This is a hands-on engineering role with significant impact on the company's next-generation AI capabilities.
Key responsibilities include: sourcing new audio data and integrating it into ingestion pipelines; operating and extending cloud infrastructure (GCP, Terraform) for data ingestion; collaborating with research scientists to optimize the cost/throughput/quality frontier; and working with leadership to shape the AI team's dataset roadmap for consumer and enterprise products.
You should have 5+ years of software development experience, strong proficiency in Python/bash scripting and Linux environments, hands-on experience with Docker and Infrastructure-as-Code, and professional experience with a major cloud provider (GCP preferred). Experience with web crawlers and large-scale data processing workflows is a plus. The ideal candidate holds a BS/MS/PhD in Computer Science or related field and can juggle multiple priorities while communicating effectively.
Speechify offers a fast-growing, entrepreneurial environment with a hands-off management approach, competitive compensation, and the opportunity to build products that directly impact people with learning differences including dyslexia, ADD, low vision, and autism.