SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Data Engineer, Data Platform (Lakehouse & Streaming)

Duetto - Remote - Remote - posted 2026-08-07

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Duetto is the hospitality industry's leading revenue management platform, trusted by hotels, resorts, and casinos worldwide. Founded in 2012 by former Wynn Resorts executives, the company has been named #1 Revenue Management Software by HotelTechAwards four years running and #1 Best Place to Work in Hotel Tech in 2025. Backed by GrowthCurve Capital since 2024, Duetto is accelerating AI investment across its product suite. You will own the design, performance, and reliability of Duetto's data lakehouse infrastructure. This is a technical authority role where you'll architect and evolve the Python/PySpark pipeline framework across a bronze → silver → gold architecture on AWS, including Glue jobs, Iceberg MERGE operations, schema evolution, and partitioning strategies. Key responsibilities include: architecting the shift from batch to near-real-time streaming with SQS-driven pipelines and Iceberg sinks; driving data quality and governance at scale through Great Expectations and data contracts; strengthening observability via Datadog, Sentry, and Sumo Logic; optimizing Glue job performance; building and maintaining shared Python libraries published to JFrog; and improving CI/CD workflows with GitHub Actions and Docker-based testing. You'll work in an AI-first engineering culture, using Claude Code and MCP tools daily, contributing to AI-assisted pipeline generation, schema inference, and automated data quality alongside a custom multi-agent system with 17 specialized agents. Required: 7+ years building production data systems in Python; deep expertise in PySpark and distributed data processing (Glue, EMR, or Databricks); strong lakehouse architecture experience (Iceberg, Delta Lake, or Hudi on S3); production Airflow or equivalent orchestrator experience; solid AWS production experience (S3, Glue, Athena, Lambda, SQS); track record improving data quality, governance, and pipeline reliability at scale. Desirable: Java knowledge for reading upstream systems; Trino or Presto experience.

Similar roles