SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Preply is a unicorn edtech platform (Series D, $150M) that connects 100,000+ tutors teaching 90+ languages to learners across 180 countries. The company is building human-led, AI-enhanced learning experiences at global scale.
As a Senior Data Engineer on the Data Ingestion and Enrichment team, you will design and own the data layer powering Preply's analytics, machine learning, and product features. You'll work across ML Platform, Data Science, Analytics Engineering, and Product teams to ensure all data assets are production-grade, observable, and reusable.
Key responsibilities:
• Design and own Preply's data lake, establishing clear ownership, schemas, and quality expectations for all datasets from ingestion through downstream consumption.
• Build and operate scalable batch and streaming ingestion pipelines supporting both real-time and analytical use cases, with explicit raw → standardized → consumption layer architecture.
• Define and implement data contracts between producers and consumers, covering schema, freshness, volume, and quality guarantees. Embed validation and anomaly detection early in the ingestion lifecycle.
• Build enrichment logic that joins, standardizes, and contextualizes data across domains, supporting historical tracking and point-in-time correctness for confident analysis.
• Instrument pipelines with strong observability (freshness, latency, quality, cost) and contribute to SLOs, alerting, and incident response playbooks.
• Apply governance and compliance by design: access control, data classification, privacy protections, masking, and audit trails embedded at ingestion time.
• Enable self-service data onboarding through standardized templates, shared libraries, and platform tooling that teams can use independently within clear guardrails.
• Collaborate cross-functionally with Product, Backend, Analytics, and ML teams to align on requirements, trade-offs, and priorities.
Required experience:
• Proven track record building architectural patterns in large, high-scale applications (APIs, high-volume pipelines, efficient algorithms).
• Platform or data engineering experience with evidence of leading multi-stakeholder deliveries.
• Cloud platform expertise (AWS/GCP) and modern DevOps practices.
• Hands-on experience designing real-time and batch data processing using Spark, Flink, Kafka, Debezium, or similar frameworks.
• Proficiency with orchestration tools (Airflow, dbt, or equivalent).
• Strong problem-solving, proactive mindset, and continuous improvement focus.
• Excellent cross-functional communication skills (English B2+).