SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
H1 is building a platform to democratize access to healthcare information globally, using data and AI to unlock medical insights and improve patient outcomes. The Data Engineering team is responsible for creating the world's most comprehensive knowledge base of healthcare stakeholders by processing thousands of data sources worldwide.
As Director of Software Engineering (Data), you will lead the Emerald team, H1's healthcare entity resolution platform responsible for linking large-scale external datasets—including PubMed, clinical trials, ClinicalTrials.gov, conferences, and web-collected data—to H1's canonical physician and organization profiles. This role sits at the intersection of distributed data engineering, entity matching, identity resolution, and large-scale healthcare data processing.
You will lead a small team of engineers while remaining deeply hands-on technically. You'll own the systems and pipelines powering automatching, grouping logic, identity mapping, deduplication, and enrichment workflows processing tens of millions of records. Key responsibilities include:
- Design, optimize, and scale distributed Spark/PySpark pipelines for entity resolution and healthcare data processing
- Own systems supporting automatching, identity mapping, deduplication, enrichment, and auto-approval workflows
- Build scalable processing frameworks for PubMed, clinical trials, ClinicalTrials.gov, conferences, and other healthcare data sources
- Drive infrastructure optimization initiatives to improve throughput, runtime, observability, and cloud compute efficiency
- Partner with AI/ML teams to integrate matching and resolution models and improve precision and recall
- Lead complex technical initiatives from architecture through deployment and production support
- Mentor engineers through code reviews, technical guidance, and best practices
- Collaborate with Product and business stakeholders to align technical solutions with operational needs
- Support production operations, incident response, and platform reliability
You bring 8+ years of experience building and maintaining large-scale distributed data systems, with deep expertise in Apache Spark, AWS (EMR, S3), and cloud-native environments. You have demonstrated technical leadership mentoring engineers and driving complex initiatives. Strong proficiency in Python (PySpark), Scala, Java, or equivalent languages is required. Experience with entity resolution, identity mapping, automatching, or large-scale matching systems is strongly preferred. You understand distributed systems fundamentals, can write production-grade code, and excel at improving performance, scalability, and observability. Familiarity with modern development tooling (Git, CI/CD, Docker, Kubernetes, Terraform, Argo, Hudi) is expected. You communicate effectively across technical and non-technical stakeholders and thrive in fast-paced, highly technical environments.