SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Data Infrastructure & Acquisition - Prague, Czech Republic

Speechify - Remote - Remote - posted 2026-08-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 30,000 - 100,000 / annual

Speechify is a text-to-speech platform used by over 50 million people globally to convert PDFs, books, Google Docs, articles, and websites into audio. The company was recently named Chrome Extension of the Year by Google and received Apple's 2025 Design Award for Inclusivity. With ~200 employees distributed globally and no physical office, Speechify operates as a fully remote organization. This role joins the Data side of Speechify's AI team, responsible for all aspects of data collection to support model training operations. The team builds high-quality datasets at petabyte scale through tight integration of infrastructure, engineering, and research. Key responsibilities include: - Sourcing new audio data sources and integrating them into the ingestion pipeline - Operating and extending cloud infrastructure for the ingestion pipeline on GCP, managed with Terraform - Collaborating with research scientists to optimize the cost/throughput/quality frontier, delivering richer datasets at scale and lower cost - Working with the AI team and leadership to define the dataset roadmap for next-generation consumer and enterprise products Required qualifications: - BS/MS/PhD in Computer Science or related field - 5+ years of industry software development experience - Proficiency with bash/Python scripting in Linux environments - Professional experience with Docker and Infrastructure-as-Code concepts - Experience with at least one major cloud provider (GCP preferred) - Preferred: web crawlers, large-scale data processing workflows - Strong communication and ability to manage multiple priorities The role offers competitive salary, equity, and the opportunity to impact a fast-growing product in the AI and audio intersection, supporting users with learning differences including dyslexia, ADD, low vision, and autism.

Similar roles