NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Data Engineer in United States of America
6 days ago
Apply with autofill
Apply with autofill
H1·6 days ago
6 days ago

Data Engineer II- Life Sciences

New York, United States of AmericaFull-timeHybridMid · 3+ yearsData Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at h1

How you compare FREE

?
Your scoreYour score: not yet known
→
67
Top 10%Top 10%: 67 out of 100

Top 10% of NextRaise users matched against Data Engineer roles in United States.

Must-have skills for this role

  • apache spark
  • pyspark
  • python
  • aws

PDF or DOCX · no account needed

Apply faster with autofill FREEh1 uses Lever - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Lead the design, optimization, and scalability of distributed Spark/PySpark pipelines powering entity resolution and large-scale healthcare data processing.
  • Own systems supporting automatching, identity mapping, grouping logic, deduplication, enrichment, and auto-approval workflows across healthcare provider and organization datasets.
  • Build and maintain scalable processing frameworks for PubMed, clinical trial, ct.gov, conference, and other healthcare data sources.
  • Drive infrastructure optimization initiatives focused on improving throughput, runtime, observability, and cloud compute cost efficiency.
  • Partner closely with AI/ML teams to integrate matching and resolution models into EMERALD and improve matching precision and recall.
  • Lead complex technical initiatives from architecture and design through deployment, monitoring, and long-term production support.
  • Serve as a technical leader and mentor across the team through code reviews, technical guidance, and engineering best practices.
  • Collaborate directly with Product and business stakeholders to align technical solutions with operational and customer needs.
  • Support production operations, incident response, troubleshooting, and ongoing platform reliability.

What they're looking for

  • 8+ years of experience building and maintaining large-scale distributed data systems and pipelines.
  • Demonstrated technical leadership experience mentoring engineers and driving complex technical initiatives.
  • Extensive experience with Apache Spark and AWS-based big data technologies including EMR, S3, and distributed compute environments.
  • Strong coding experience in Python (PySpark), Scala, Java, or equivalent languages used for distributed processing systems.
  • Experience optimizing large-scale Spark workloads for performance, scalability, and infrastructure cost efficiency.
  • Experience with streaming and event-driven architectures using technologies such as Kafka or Spark Streaming.
  • Experience with orchestration and lakehouse technologies such as Argo and Hudi or comparable platforms.
  • Experience with containerization and infrastructure technologies such as Docker, Kubernetes, and Terraform.
  • Experience working with relational or distributed databases such as PostgreSQL or Redshift.
  • Proven ability to operate effectively within highly scalable, production-grade distributed systems.

Nice to have

  • Experience working with healthcare, life sciences, Real World Evidence (RWE), or large-scale healthcare datasets is strongly preferred.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

At H1, we believe access to the best healthcare information is a basic human right. Our mission is to provide a platform that can optimally inform every doctor interaction globally. This promotes health equity and builds needed trust in healthcare systems. To accomplish this, our teams harness the power of data and AI-technology to unlock groundbreaking medical insights and convert those insights into action that result in optimal patient outcomes and accelerates an equitable and inclusive drug development lifecycle. Visit h1.com to learn more about us.

As part of H1’s hiring process, all candidates are required to participate in an in-person final interview. Depending on your location, this may require travel.

H1's Data Network (H1DN) team is the client-data mastering network at the core of how H1's products get their data. We run production ingestion for major enterprise customers. Clinical trial data is one of our highest-visibility streams: it feeds decisions about where trials run and who runs them, and the people who depend on it are as often clinical experts as they are engineers. SLAs and customer expectations drive how we work, and we're looking for engineers who are energized by that.

 
WHAT YOU'LL DO AT H1
As a Data Engineer II on the H1DN team, you will build and operate the pipelines behind H1's clinical trials data. You'll work primarily in Python, PySpark, and SQL, and you'll work directly with clinical subject matter experts and Customer Success Managers to turn their domain knowledge into pipeline logic that holds up in production.

You will:
- Build and maintain the Python and PySpark pipelines behind the CTMS trial data pipeline intake workflows, including scoring and status logic.
- Develop the transformation logic that maps raw trial and customer data to H1's internal data models, handling diverse source formats including CSV, JSON, Parquet, and APIs.
- Write and tune SQL against large datasets to investigate data questions, validate pipeline output, and support analysis that clinical SMEs and customer-facing teams depend on.
- Turn around customer-driven changes quickly, scoping requests as they arrive, shipping changes that hold up under enterprise SLAs, and reworking logic as customer needs shift mid-flight.
- Partner with clinical SMEs to translate domain expertise into concrete data rules, then walk them through the results, explain what the pipeline did and why, and fold their feedback back into the logic.
- Build the data quality checks, validation logic, and reconciliation that let non-engineers trust pipeline output without reading the code.
- Participate in code reviews, maintaining a high bar for quality and adherence to engineering standards.
- Monitor and improve pipeline observability, contributing to alerting and dashboards that surface job health and data anomalies for both the team and internal users.
 
ABOUT YOU
You are a data engineer with a strong Python foundation and real distributed-processing experience. You're drawn to high-impact teams where the work is tangible: pipelines running, enterprise customers getting their data on time, clinical data that people make real decisions from. You're comfortable in an environment where recurring production runs and customer SLAs shape day-to-day priorities, and where a customer request can reorder your week. You'd rather sit down with a domain expert and understand why the data looks the way it does than build to a spec handed to you secondhand.
 
You bring experience:
- Building and shipping production data pipelines in Python, with an understanding of what makes them reliable and maintainable under real load
- Working with PySpark or a comparable distributed processing framework on datasets too large for a single machine
- Writing SQL well enough to answer hard questions about data, not just retrieve it
- Working in an operationally-driven environment where reliability and on-time delivery matter as much as new feature work
- Working directly with non-engineering partners, subject matter experts, analysts, or customer-facing teams, and communicating clearly about data with people who don't read code
- Holding a high bar in code review and expecting the same from those who review your work
- Identifying data problems early and seeing work through to resolution rather than handing it off
 
REQUIREMENTS 
- 3+ years of experience in software or data engineering, with meaningful Python in your background
- Demonstrated experience building and maintaining production-grade data pipelines in Python
- Hands-on experience with PySpark or a similar distributed data processing framework
- Strong SQL skills, including working with large, messy, multi-source datasets
- Strong understanding of software quality practices: testing, code review, documentation, and CI/CD
- Experience working with cross-functional and non-technical stakeholders
- Experience with pipeline orchestration tooling (Argo, Airflow, Databricks, dbt, or similar) preferred
- Familiarity with clinical trial data, healthcare data, or another regulated data domain a plus
- Familiarity with entity matching or data mastering a plus
- Familiarity with AWS services (S3, Lambda, ECS, or similar) a plus
 
 
COMPENSATION
This role pays $110,000 to $135,000 per year, based on experience, in addition to stock options.

Anticipated role close date: 10/20/2026

Company

H1
New York, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from H1's careers site·first seen 15 Sept 2026·last verified 15 Sept 2026·How we source jobs

Similar jobs

  • NCIS Data Engineer | Active Secret clearance at gditUSA VA Quantico–match not yet calculated
  • Senior Data Engineer, Selling Partner Insights and Analytics at AmazonSeattle, United States of America–match not yet calculated
  • Senior Data Engineer (Scala) - Remote (USA) at icfReston, United States of America–match not yet calculated
  • Analytics Engineer at fireworksSan Mateo, United States of America–match not yet calculated
  • Data Engineer at michelinhrGREENVILLE, United States of America–match not yet calculated

Browse more jobs

  • Data Engineer jobs in United States
  • Data Scientist jobs in United States
  • Data Analyst jobs in United States
  • Business Intelligence Analyst jobs in United States
  • Data Engineer jobs in India
  • Data Engineer jobs in United Kingdom