Skip to content

Reference architecture

HL7 and FHIR healthcare data lakehouse on Databricks

Unify fragmented EHR feeds into a FHIR R4 lakehouse on Delta Lake, with de-identified datasets for research and access controls built for HIPAA.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Ingest HL7 v2, C-CDA and FHIR R4 records from Epic and Oracle Health (Cerner) EHRs
  • 02Automated de-identification using the HIPAA Safe Harbor method
  • 03Data available for research and population health queries within a minute
  • 04Full lineage from raw feeds to analytics marts

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

HL7 and FHIR healthcare data lakehouse on Databricks

Illustrative reference architecture

  1. 01

    HL7/FHIR ingestion gateway

    Receives MLLP HL7 feeds and RESTful FHIR bundles

    AWS HealthLake / Apache Camel

  2. 02

    Medallion data lakehouse

    Bronze (raw), silver (FHIR-normalized) and gold (clinical marts) layers

    Delta Lake + Databricks

  3. 03

    De-identification engine

    Removes the 18 HIPAA Safe Harbor identifiers for research access

    John Snow Labs NLP / Amazon Comprehend Medical

  4. 04

    Clinical analytics engine

    SQL for clinical trials and cohort analysis

    Databricks SQL / Trino

Subsystem 01

HL7/FHIR ingestion gateway

Receives MLLP HL7 feeds and RESTful FHIR bundles

Typical stack

AWS HealthLake / Apache Camel

Subsystem 02

Medallion data lakehouse

Bronze (raw), silver (FHIR-normalized) and gold (clinical marts) layers

Typical stack

Delta Lake + Databricks

Subsystem 03

De-identification engine

Removes the 18 HIPAA Safe Harbor identifiers for research access

Typical stack

John Snow Labs NLP / Amazon Comprehend Medical

Subsystem 04

Clinical analytics engine

SQL for clinical trials and cohort analysis

Typical stack

Databricks SQL / Trino

Data lifecycle

End-to-end data flow

  1. Hospital EHRs send real-time HL7 v2 ADT and ORU messages over an encrypted VPN to the ingestion gateway.

  2. The gateway converts messages into FHIR R4 JSON resources and writes them to bronze Delta tables.

  3. A Spark pipeline validates schemas, resolves duplicate patient IDs and writes silver normalized tables.

  4. The de-identification pipeline strips the Safe Harbor identifiers to produce gold research datasets.

  5. Researchers query population health datasets with standard SQL and Tableau dashboards.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Duplicate patient identities across hospital feeds

Mitigation

A probabilistic master patient index (MPI) using Fellegi–Sunter record linkage.

Failure mode 02

Malformed custom HL7 segments

Mitigation

Quarantine invalid messages in a dead-letter S3 bucket and alert data engineering in Slack.

Failure mode 03

Research access to identified ePHI

Mitigation

Unity Catalog role-based access control with dynamic column masking.

Questions

What teams ask about this design

FHIR (Fast Healthcare Interoperability Resources) is the HL7 standard for exchanging healthcare data through APIs. In the US, ONC certification and CMS interoperability rules require FHIR R4 APIs from certified EHRs and many payers.

It organizes data into bronze (raw ingested data), silver (cleaned and normalized) and gold (aggregated, business-level clinical data) layers.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.