SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer — Distributed Compute / Spark Systems

Granica - Mountain View, CA, United States - In-office - posted 2026-08-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Granica builds AI infrastructure for enterprises managing massive data environments. The company's platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI. Products include Crunch (continuous optimization for enterprise lakehouse data), Myelin (stateful infrastructure for long-running AI agents), and Large Tabular Models (foundation models for enterprise tables). You will join as a Senior Software Engineer focused on building distributed compute systems for enterprise-scale data and AI workloads. You will own core infrastructure behind Crunch, including systems for distributed execution, workload optimization, query performance, scheduling, resource management, and compute cost reduction across petabyte- and exabyte-scale environments. This is a hands-on role where you will directly impact customer compute spend, query latency, workload reliability, cluster efficiency, and large-scale analytical data processing performance. Key responsibilities include: building distributed compute systems for large-scale analytical and AI workloads; improving performance and cost efficiency across Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments; designing workload-aware systems for query execution, resource allocation, and scheduling; optimizing execution performance across joins, aggregations, scans, shuffles, spills, caching, partitioning, and task scheduling; building systems that learn from workload patterns to automatically improve execution plans and cluster usage; developing infrastructure for adaptive workload routing and data-processing reliability; debugging performance bottlenecks across query execution, metadata, storage, network, memory, CPU, and distributed compute layers; working with lakehouse tables and columnar formats (Iceberg, Delta Lake, Hudi, Parquet, ORC); building systems to reduce compute waste from inefficient scans, poor partitioning, small files, skew, and suboptimal workload placement; improving reliability and failure recovery for large distributed data-processing jobs; implementing algorithms in workload optimization and cost modeling; and contributing to open-source projects or publishing research when appropriate. You will partner directly with Product, Engineering, and company leadership, helping shape Crunch and influencing architecture, product direction, customer outcomes, and company growth. REQUIREMENTS: - Strong engineering depth in distributed systems, data processing systems, query engines, databases, or cloud infrastructure - Production experience with distributed compute or query systems such as Apache Spark, Spark SQL, Trino, Presto, Flink, Databricks, EMR, Glue, Hive, or similar systems - Hands-on experience improving performance, reliability, or cost efficiency for large-scale data-processing workloads - Understanding of distributed execution, query planning, scheduling, resource management, fault tolerance, and workload isolation - Experience with Spark internals, Spark SQL, Catalyst, Adaptive Query Execution, shuffle, joins, aggregation, spill, memory management, or task scheduling - Familiarity with lakehouse formats and columnar data such as Iceberg, Delta Lake, Hudi, Parquet, or ORC - Familiarity with cloud object storage systems (S3, GCS, ADLS) and performance tradeoffs of running distributed compute on top of them - Strong programming skills in Scala, Java, Go, Rust, C++, or similar systems-oriented languages - Curiosity about workload optimization, cost modeling, adaptive execution, and how compute efficiency affects AI and analytics at scale - Pragmatic builder's mindset: rigorous, hands-on, and comfortable owning complex systems end to end BONUS QUALIFICATIONS: - Experience contributing to Apache Spark, Spark SQL, Trino, Presto, Flink, Velox, DuckDB, DataFusion, Iceberg, Delta Lake, Hudi, Parquet, ORC, or related systems - Experience with Catalyst, Adaptive Query Execution, cost-based optimization, query planning, vectorized execution, or distributed runtime systems - Experience optimizing joins, aggregations, shuffles, scans, spills, caching, partitioning, skew handling, or task scheduling - Experience building workload schedulers, execution control planes, resource managers, or multi-engine compute platforms - Experience reducing compute cost or improving workload efficiency in large-scale production data environments - Background in query engines, distributed runtimes, storage-aware execution, indexing, caching, encoding, compression, or adaptive query optimization - Research or open-source contributions in distributed systems, databases, query processing, data processing, or cloud infrastructure

Similar roles