SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 175,000 - 245,000 / annual
Superhuman (formerly Grammarly, now part of the Superhuman AI productivity platform) is seeking a Data Engineer to join the Data Foundations team within the Data Platform organization. This role focuses on building world-class data infrastructure to support Superhuman's mission of unlocking superhuman potential through AI-powered productivity tools.
The Data Foundations team owns foundational datasets and data models that define core business entities across Superhuman's integrated suite of products (Grammarly writing assistance, Docs collaborative workspace, Mail inbox management, and Go proactive AI assistant). The team processes over 70 billion daily events and manages petabyte-scale data infrastructure.
In this role, you will:
- Design and implement large-scale, robust data pipelines and data lakes using Spark/Databricks that ingest and process billions of daily events into reliable, decision-grade datasets
- Architect solutions for data availability, security, and scalability across the platform, supporting both real-time and batch processing
- Own data quality, freshness, and reliability for foundational datasets, implementing automated checks, monitoring, and alerting
- Build ETL frameworks and tooling that enable other Data teams and data scientists to self-serve and model core business entities into clean, well-documented, reusable datasets
- Partner with product teams, backend engineers, ML engineers, and Data Science teams to translate business questions into high-impact data products
- Collaborate with leadership to shape the team's technical roadmap and strategy
This is a high-impact role at the intersection of data engineering, data-intensive applications, and data modeling. The systems you build directly enable other data teams' efficiency and support business-critical features, research, and experimentation across the company's integrated product suite.
The role offers a hybrid working model with preferred locations in San Francisco or Seattle hubs, providing a balance of focus time and in-person collaboration.
QUALIFICATIONS (Required):
- 3+ years of experience running live production environments, including high-load or data-intensive workflows, with focus on uptime and reliability
- Proficiency in SQL and Python, with deep hands-on experience in Spark and a modern lakehouse or cloud data warehouse (Databricks, Delta Lake, dbt, Snowflake, or similar)
- Strong knowledge of ETL/ELT design patterns, orchestration tools (Airflow, dbt, Dagster, or similar), data quality frameworks, and CI/CD for data with Git-based deployments
- Data modeling and data warehouse design skills with a rigorous approach to data quality and observability
- Hands-on experience with modern data storage technologies (Delta Lake, Snowflake, BigQuery, Redshift, or similar)
- Clear communication and strong collaboration skills across diverse teams and stakeholders
- Strong ownership mindset with end-to-end responsibility from strategy and architecture through production and ongoing reliability
- Strong acumen for scalable, highly complex data processing with ability to think holistically and raise the bar on team craft
- Self-starting problem-solver who thinks from first principles, manages multiple priorities, and thrives in fast-paced, results-driven environments
NICE TO HAVE:
- Experience operating large-scale distributed systems in data engineering and data modeling
- Experience with high-throughput, real-time streaming systems (Kafka, Flink, Spark Structured Streaming) at billions-of-events-per-day scale
- Comfort with data lake and lakehouse technologies (Delta Lake, Iceberg, Hudi) and cloud infrastructure as code (Terraform)
- Track record of building self-serve data products, tools, or frameworks that other teams rely on
- Experience partnering with analytics, data science, or ML teams as a strategic data partner to productionize data and models