SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
inDrive is seeking a Data Engineer to join one of its Data Platform teams, working across Marketing, Growth, partner, and financial data domains. You will design and operate large-scale data infrastructure using cutting-edge cloud technologies including GCP, AWS, BigQuery, Databricks, and Kubernetes.
Key responsibilities include:
• Build and operate batch and streaming ingestion pipelines into a layered BigQuery data warehouse (raw → ODS → data marts) using Airflow, Debezium CDC over Kafka with protobuf, Pub/Sub, and Dataflow
• Integrate external data sources end-to-end, including marketing platforms (GA4, AppsFlyer, TikTok/Meta/Google Ads), payment providers, S3 buckets, and third-party APIs, with schema contracts, backfills, and reconciliation
• Engineer the data platform in Python: develop custom Airflow operators and connectors in a shared ETL framework, Kafka Connect on Strimzi (Kubernetes), Cloud Functions, and API integrations
• Build CI/CD and change-management tooling for BigQuery: GitHub-based test-and-approval flows, SQL migration engines (Liquibase/Flyway/Bytebase), sandbox validation, backup and rollback
• Own pipeline reliability and correctness: implement idempotency, deduplication, late-data handling, backfill and replay logic, freshness monitoring and alerting; write integration and unit tests
• Drive data governance and compliance: ITGC-compliant change management for BigQuery, IAM and least-privilege access, PII policy tags and DLP, Unity Catalog on Databricks, column-level lineage (OpenMetadata/Dataplex), and disaster-recovery planning
• Build internal data tools and platform services for agentic workflows: Streamlit apps, Slack bots, LLM-based agents and MCP servers to help teams find and use data
• Support analysts and business teams with data requests and foster data-driven decision-making
• Contribute to system design and architecture with the development team
REQUIREMENTS:
• Strong practical Python: clean, well-structured, and tested code for services, tooling, and data pipelines
• Solid software design skills (OOP, modularity, design patterns)
• Experience building and operating services in a cloud environment (GCP, AWS, or similar): CI/CD, containerization, monitoring and alerting
• Familiarity with Kubernetes and Terraform
• Hands-on experience with data warehouse tasks (BigQuery or another cloud warehouse) and confident SQL skills
• Clear communication with non-engineering stakeholders
• Demonstrated ability to take ownership of technologies or services and proactively contribute ideas
NICE TO HAVE:
• Advanced SQL: complex queries, window functions, partitioning, clustering, and cost optimization
• Experience building reliable CDC pipelines (e.g., Debezium): idempotency, schema evolution, backfills, and reconciliation
• Analytical data modeling skills: table grain, facts vs dimensions, slowly changing dimensions, and metric definitions
• Experience with stream processing frameworks such as Flink or Apache Beam/Dataflow
• Exposure to data governance and audit compliance (ITGC/SOX), Databricks Unity Catalog, or lineage/catalog tooling
• Interest in building LLM-based agents and AI tooling for data