SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Speechify is a 50M+ user text-to-speech platform that transforms reading into audio, enabling faster comprehension and retention. The company has been recognized by Google (Chrome Extension of the Year) and Apple (2025 Design Award for Inclusivity). Operating as a fully distributed 200-person team with talent from Amazon, Microsoft, Google, Stanford, and high-growth startups, Speechify is building next-generation AI-powered audio products.
You'll join the Data side of the AI team, responsible for all aspects of data collection supporting model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research.
Key responsibilities:
- Source and integrate new audio data streams into the ingestion pipeline
- Operate and extend cloud infrastructure (GCP, Terraform) for data ingestion
- Partner with research scientists to optimize the cost/throughput/quality frontier, delivering richer datasets at scale and lower cost
- Collaborate with the AI team and leadership to define the dataset roadmap powering next-generation consumer and enterprise products
Required qualifications:
- BS/MS/PhD in Computer Science or related field
- 5+ years of industry software development experience
- Proficiency with bash/Python scripting in Linux environments
- Professional experience with Docker and Infrastructure-as-Code (GCP preferred)
- Strong communication skills
Desirable experience includes web crawlers and large-scale data processing workflows. The role offers a fast-growing environment, entrepreneurial culture, hands-off management, competitive compensation, and the opportunity to impact millions of users with accessibility-focused technology.