SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Preply is a unicorn edtech platform (Series D, $150M) that connects 100,000+ tutors teaching 90+ languages to learners across 180 countries. The Data Ingestion and Enrichment team builds the trusted, scalable data foundation powering analytics, machine learning, and product features.
As a Data Engineer, you will design and operate the data layer that serves analytics, ML, and product teams. You'll build end-to-end batch and streaming ingestion pipelines, implement data quality contracts and validation, and ensure every dataset has clear ownership, schema, and quality expectations. Key responsibilities include:
• Build and maintain components of Preply's data lake, treating trust, correctness, and predictability as first-class features.
• Develop reliable batch and streaming pipelines (Spark, Flink, Kafka, Debezium) that support real-time and analytical use cases.
• Implement data contracts covering schema, freshness, volume, and quality guarantees; embed validation and anomaly detection early in the ingestion lifecycle.
• Build enrichment logic that joins, standardizes, and contextualizes data across domains; support historical tracking and point-in-time correctness.
• Instrument pipelines with strong observability (freshness, latency, quality, cost); contribute to SLOs, alerting, and incident response.
• Apply governance, access control, and privacy protections at ingestion time; ensure sensitive data is masked or anonymized by default.
• Enable self-service through standardized templates, shared libraries, and improved discoverability so teams can onboard data sources independently.
• Collaborate closely with Product, Backend, Analytics, and ML teams; mentor junior engineers and foster a culture of data quality standards.
You'll need hands-on experience building large-scale applications, platform or data engineering background, cloud platform familiarity (AWS/GCP), and expertise with modern data frameworks. Experience with orchestration tools (Airflow, dbt) and strong cross-functional communication (English B2+) are essential.