SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Development Engineer II – Data Engineer

Gather AI - Remote - Remote - posted 2026-09-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Gather AI is building a vision-powered platform using autonomous drones and existing equipment to digitize warehouse operations and supply chain workflows. The company is pioneering warehouse intelligence by capturing real-time data to replace manual, error-prone processes. You'll join the Full Stack group within Cloud Services as one of the first engineers building Gather AI's data foundation from scratch. Today, production and analytical workloads share a single database, and each product defines its own metrics. This team is designing and building the warehouse, transformation layers, and semantic model from the ground up, working closely with Full Stack, ML, Product, Security, and Customer Success teams. Unlike typical data engineering roles that maintain one slice of a mature platform or build against pre-designed models, you'll have the rare opportunity to architect foundational data infrastructure. You'll start by shipping well-scoped pipelines and models on the Drone product alongside senior engineers, then grow toward owning a data domain end to end: its design, quality, and on-call responsibilities. Key responsibilities include: - Building extraction pipelines to move data from production PostgreSQL into the analytical warehouse using incremental loads - Writing and maintaining dbt models (dimensions, reusable metric building blocks, serving tables) following the Principal Data Engineer's architecture - Implementing metrics in the semantic layer so measures are defined once and consistent across all dashboards - Adding tests, freshness checks, and alerts to ensure data trustworthiness - Applying tenant isolation and access-control patterns to every model and pipeline - Maintaining lineage so metrics are traceable to source records and drone images - Partnering with the integration team to validate incoming WMS data - Documenting work and registering models in the data catalog - Delivering through CI/CD with AI-assisted development and joining on-call for owned pipelines - Growing the platform by onboarding MHE Vision, 3D case counting, and other products onto the shared data foundation This is a fully remote role on the India-based team. Clear written communication, initiative, and the ability to make progress independently are essential given the distributed, multi-timezone environment. REQUIREMENTS: - 2–5 years building and running production data pipelines; degree in Computer Science or equivalent practical experience - Strong SQL (joins, window functions, CTEs) with working understanding of dimensional modeling (facts, dimensions, grain); PostgreSQL experience preferred - Hands-on experience building tested models in dbt, Snowflake Dynamic Tables, Databricks Lakeflow Declarative Pipelines, or equivalent - Production pipeline code in Python with tests and code review (beyond notebooks or one-off scripts) - Experience running pipelines in Airflow, Dagster, Databricks Lakeflow Jobs, Snowflake Tasks, or equivalent, including handling failures, re-runs, and backfills - Experience loading data from operational databases into warehouses or lakehouses (Snowflake, Databricks) - Production experience on Azure or another major cloud, including object storage, Git, CI/CD, and Docker/Kubernetes - Data quality mindset: writes tests alongside code and cares whether data is correct - Ownership and growth orientation: hands-on, curious, digs into problems beyond own code, brings ideas to team, takes feedback well - Clear written and spoken English; writes good PRs and documentation; raises blockers early in distributed teams NICE TO HAVE: - Change data capture (Debezium, Fivetran) or streaming (Kafka, Event Hubs) - PySpark or Snowpark - Terraform or infrastructure as code - Semantic layers (dbt Semantic Layer, Cube) or data catalogs (Purview, DataHub) - Multi-tenant data platforms or row-level security - Working with image, video, or sensor data alongside structured records - Logistics, warehousing, or robotics domain experience

Similar roles