SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 190,000 - 220,000 / annual
H1 is building a platform to provide optimal healthcare information globally, promoting health equity and accelerating drug development. The Data Engineering team is responsible for developing and delivering H1's most critical asset—data—by processing thousands of global healthcare data sources to create a comprehensive knowledge base of healthcare stakeholders and ecosystem relationships.
As a Staff Data Engineer on the Emerald team, you will shape the architecture, scalability, and technical direction of H1's healthcare entity resolution platform. Emerald links large-scale external healthcare datasets (PubMed, clinical trials, ct.gov, conferences, web-collected data) to H1's canonical physician and organization profiles. This role sits at the intersection of distributed data engineering, entity matching, identity resolution, and large-scale healthcare data processing.
You will lead a small team of engineers while remaining deeply hands-on technically, owning systems and pipelines powering automatching, grouping logic, identity mapping, deduplication, and enrichment workflows processing tens of millions of records. You will partner closely with Product, AI/ML, Analytics, and Engineering teams to improve platform accuracy, scalability, reliability, and operational efficiency.
Key responsibilities include:
- Design, optimize, and scale distributed Spark/PySpark pipelines for entity resolution and large-scale healthcare data processing
- Own systems supporting automatching, identity mapping, grouping logic, deduplication, enrichment, and auto-approval workflows
- Build and maintain scalable processing frameworks for PubMed, clinical trials, ct.gov, conference, and other healthcare data sources
- Drive infrastructure optimization initiatives for throughput, runtime, observability, and cloud compute cost efficiency
- Partner with AI/ML teams to integrate matching and resolution models and improve matching precision and recall
- Lead complex technical initiatives from architecture through deployment, monitoring, and production support
- Mentor engineers through code reviews, technical guidance, and engineering best practices
- Collaborate with Product and business stakeholders to align technical solutions with operational needs
- Support production operations, incident response, troubleshooting, and platform reliability
You are an experienced data engineer with deep expertise building and optimizing distributed data systems in cloud-native environments. You thrive solving complex scalability and performance challenges and enjoy operating in highly technical, fast-paced engineering environments. You bring strong hands-on engineering expertise while also guiding technical direction and mentoring engineers.
REQUIREMENTS:
- 8+ years of experience building and maintaining large-scale distributed data systems and pipelines
- Demonstrated technical leadership experience mentoring engineers and driving complex technical initiatives
- Extensive experience with Apache Spark and AWS-based big data technologies (EMR, S3, distributed compute)
- Strong coding experience in Python (PySpark), Scala, Java, or equivalent languages for distributed processing
- Experience optimizing large-scale Spark workloads for performance, scalability, and infrastructure cost efficiency
- Experience with streaming and event-driven architectures (Kafka, Spark Streaming)
- Experience with orchestration and lakehouse technologies (Argo, Hudi, or comparable platforms)
- Experience with containerization and infrastructure technologies (Docker, Kubernetes, Terraform)
- Experience with relational or distributed databases (PostgreSQL, Redshift)
- Proven ability to operate effectively within highly scalable, production-grade distributed systems
- Healthcare, life sciences, Real World Evidence (RWE), or large-scale healthcare datasets experience strongly preferred
Additional desired expertise: distributed file formats (Parquet, AVRO), entity resolution/identity mapping/automatching/deduplication systems, root cause analysis across large-scale systems, modern development tooling (Git, CI/CD, JIRA).