SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Commure is building an AI Operating System for healthcare, delivering ambient AI documentation, intelligent workflow automation, and autonomous revenue cycle management across 60+ EHR integrations. The platform processes $25B+ in annual claims and supports 200+ million patient interactions for 500+ healthcare organizations.
You will own the data warehouse platform end-to-end, designing and operating every layer of the stack: CDC pipelines, streaming transport, schema governance, query and serving layer, and analytics platform. This is a hands-on individual contributor role with broad architectural scope. You'll make critical design decisions, write code that sets patterns for other teams, and directly impact how healthcare data flows through the organization.
Key responsibilities:
- Design and operate CDC pipelines using Debezium and streaming infrastructure (Kafka, Redpanda) to move data from operational databases into the warehouse with low latency and high fidelity
- Architect the data lake on object storage using open table formats (Iceberg, Delta Lake, Hudi) with Parquet, supporting both batch and streaming workloads
- Run and scale StarRocks (or similar MPP/lakehouse engines) as the query and serving layer, including schema design, materialized views, ingestion patterns, and performance tuning
- Build the transformation layer with dbt: modeling standards, tests, documentation, and semantic layer for metrics
- Stand up orchestration (Airflow, Dagster) and CI/CD, observability, and data-quality tooling
- Partner with Security and Compliance on PHI/PII handling, access controls, lineage, and auditability to meet HIPAA and SOC 2 requirements
- Establish patterns and conventions that enable product and analytics teams to self-serve on the platform
The core stack includes Debezium for CDC, StarRocks for query and serving, and dbt for transformation. You'll work directly with clinicians and teams, deploy daily, and own work from conception to production.
REQUIREMENTS:
- 6+ years of software engineering experience with significant time building or operating data platforms at scale
- Experience across the modern data stack: CDC (Debezium or equivalent), streaming (Kafka, Redpanda), data lake formats (Iceberg, Delta, Hudi), MPP or lakehouse query engine (StarRocks, ClickHouse, Trino, Snowflake, Databricks), and dbt
- Fluent in SQL, schema design, query optimization, and reasoning about cost and latency trade-offs on large datasets
- Experience running production data infrastructure: orchestration, observability, on-call, data quality, and incident response
PREFERRED:
- Direct production experience with Debezium, StarRocks, and dbt
- Experience building semantic layers (dbt Semantic Layer, Cube) or data catalogs/lineage tools (DataHub, OpenMetadata, Amundsen)
- Experience with HIPAA-regulated data: PHI handling, de-identification, access governance
- Experience powering AI/ML workloads: feature stores, training-set curation, embedding pipelines, retrieval systems
- Multi-cloud experience (AWS, GCP, Azure), infrastructure-as-code (Terraform, Pulumi), Kubernetes controllers