SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Apna is seeking a Senior Data Engineer to design, build, and operate critical components of their data platform. This role focuses on large-scale data pipelines, lakehouse architecture, query optimization, and workflow orchestration that power analytics, product intelligence, machine learning, and business dashboards across the company.
Key responsibilities include building scalable batch and near-real-time data pipelines for product, business, growth, and ML use cases; designing and improving lakehouse architecture using Apache Hudi; working with distributed query engines like Presto/Trino for analytical workloads; building and maintaining orchestration workflows with Apache Airflow; creating reusable data models and curated datasets; improving platform reliability, observability, and data quality; optimizing storage and compute costs; and partnering with product, analytics, ML, and backend teams to translate data needs into scalable solutions.
The role also involves driving engineering standards around data modeling, schema evolution, partitioning, and pipeline ownership, as well as mentoring junior data engineers and influencing architecture decisions.
Required qualifications include 4-6 years of data engineering experience at scale, hands-on expertise with Apache Airflow or similar orchestration systems, strong knowledge of Presto/Trino and Apache Hudi concepts (copy-on-write, merge-on-read, upserts, compaction, clustering, timeline management, schema evolution), deep understanding of distributed data processing and storage systems, ability to design reliable ETL/ELT pipelines, strong SQL skills, familiarity with various data architectures (data warehouse, data lake, lakehouse, lambda, kappa, medallion, event-driven), experience with data modeling for analytics, and strong programming skills in Python, Java, or Scala.
Desirable experience includes Kafka, Spark, Flink, Hive, Iceberg, Delta Lake, BigQuery, internal data platform development, data quality frameworks, ML feature pipelines, metadata management and data catalogs, cloud infrastructure (AWS, GCP, Azure), and privacy/compliance knowledge.