SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Apna is seeking a Senior Data Engineer to design, build, and operate critical components of their data platform. You will work on large-scale data pipelines, lakehouse architecture, query platforms, and workflow orchestration systems that power analytics, product intelligence, machine learning, business dashboards, experimentation, and operational decision-making across the company.
Key responsibilities include building scalable batch and near-real-time data pipelines across product, business, growth, and ML use cases; designing and improving lakehouse architecture using Apache Hudi; working with query engines like Presto/Trino for analytical workloads; building and maintaining orchestration workflows using Apache Airflow; creating reusable data models and curated datasets for analytics and product teams; improving platform reliability, observability, SLA tracking, lineage, and data quality checks; optimizing storage, compute, query performance, and pipeline costs; partnering with product, analytics, ML, and backend teams to understand data needs and convert them into scalable solutions; driving engineering standards around data modeling, schema evolution, partitioning, deduplication, and pipeline ownership; and mentoring junior data engineers.
Required qualifications include strong experience in data engineering at scale; hands-on experience with Apache Airflow or similar orchestration systems; strong knowledge of Presto/Trino or other distributed query engines; deep understanding of Apache Hudi concepts including copy-on-write vs merge-on-read, upserts, deletes, incremental reads, compaction, clustering, timeline/commits, schema evolution, and partitioning; strong knowledge of distributed data processing and storage systems; ability to design and build reliable ETL/ELT pipelines; strong SQL skills and ability to debug complex data issues; understanding of various data architectures (data warehouse, data lake, lakehouse, lambda, kappa, medallion, event-driven); experience with data modeling for analytics and reporting; strong programming skills in Python, Java, or Scala; and ability to reason about trade-offs between freshness, cost, reliability, latency, and complexity.
Desirable experience includes Kafka, Spark, Flink, Hive, Iceberg, Delta Lake, or BigQuery; building internal data platforms or self-serve infrastructure; data quality frameworks like Great Expectations or Deequ; ML feature pipelines or feature stores; metadata management and data catalogs; cloud infrastructure (AWS, GCP, Azure); and privacy/compliance/PII handling in data systems. The role requires 4-6 years of experience and is based in Domlur, Bangalore with 5 days per week in-office work.