SlipstreamJobsFresh Startup & VC-Backed Jobs

Principal Software Engineer - Data Platform (Iceberg/Trino)

Innovaccer - United States - In-office - posted 2026-08-13

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Innovaccer is seeking a Principal Software Engineer to architect and own the lakehouse data platform that powers on-premise deployments. This role is responsible for designing a ground-up lakehouse engine built on Apache Iceberg (table format), Trino (query engine), a REST catalog service, and Spark (transform compute)—a critical technical track with decisions that gate downstream transforms, serving, and reporting. Key responsibilities include: - Owning the complete lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark compute, and object-storage layout - Designing on-premise equivalents for cloud-managed warehouse capabilities (CDC streams, scheduled tasks, write-back paths) - Running proof-of-concept validation at expected data volumes and defining evidence-based triggers for placement decisions (VM vs. Kubernetes) - Setting platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance (compaction, snapshot expiry, orphan cleanup) - Leading SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL - Mentoring senior engineers across data workstreams, reviewing designs, and raising engineering quality standards - Partnering with platform engineering on storage sizing, resource isolation, and capacity planning Required qualifications: - B.E., B.Tech., or M.Sc. in Computer Science or related technical field - 12+ years building and operating large-scale data platforms or distributed systems - Deep hands-on expertise with distributed SQL engines (Trino/Presto or Spark SQL): query planning, performance engineering, internals - Production experience with Apache Iceberg (or Delta Lake/Hudi with willingness to specialize in Iceberg): table spec, merge-on-read vs. copy-on-write, maintenance at scale - Working knowledge of Iceberg catalog services (REST catalogs like Polaris or Nessie, or Hive Metastore) and S3-compatible object storage - Strong understanding of cloud warehouse internals (Snowflake, BigQuery, Redshift) to design functional equivalents on open-source infrastructure - Professional software development in Java and/or Python - Experience with on-premise, regulated, or air-gapped environments is a strong plus; healthcare data experience is a plus Innovaccer offers competitive benefits including 20 days PTO, generous parental leave, recognition programs, and comprehensive insurance (medical, dental, vision, disability, life).

Similar roles