SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Preply is a unicorn edtech platform (Series D, $150M) that connects 100,000+ tutors teaching 90+ languages to learners across 180 countries. The Data Ingestion and Enrichment team builds the trusted data foundation powering analytics, machine learning, and product features.
As a Data Engineer, you will own critical components of Preply's data lake and ingestion infrastructure. Your responsibilities include:
**Core Responsibilities:**
- Build and maintain scalable batch and streaming ingestion pipelines that support both real-time and analytical use cases, balancing performance, cost, and reliability.
- Design and implement data contracts between producers and consumers, covering schema, freshness, volume, and quality guarantees.
- Embed validation, anomaly detection, and quality checks early in the ingestion lifecycle to prevent issues from propagating downstream.
- Develop enrichment logic that joins, standardizes, and contextualizes data across domains using shared definitions and reusable patterns.
- Support historical tracking, point-in-time correctness, and dataset versioning for confident downstream analysis.
- Instrument pipelines with strong observability (freshness, latency, quality, cost metrics) and contribute to SLOs and incident response playbooks.
- Apply consistent access control, classification, and privacy protections at ingestion time, ensuring sensitive data is properly masked or anonymized.
- Contribute to standardized ingestion templates and platform tooling that enable self-service data onboarding.
- Collaborate closely with Product, Backend, Analytics, and ML teams to align on requirements and trade-offs.
**Required Experience:**
- Hands-on experience building large-scale applications (data pipelines, APIs, algorithms).
- Platform or data engineering team experience in multi-stakeholder environments.
- Cloud platform proficiency (AWS/GCP) and modern DevOps practices.
- Real-time and batch data processing with Spark, Flink, Kafka, Debezium, or similar frameworks.
- Orchestration tools experience (Airflow, dbt, or equivalent).
- Strong problem-solving and proactive innovation mindset.
- Excellent cross-functional communication skills (English B2+).
Preply offers competitive compensation with equity, learning allowances, mental health support, and the opportunity to impact education globally.