Database Credentialed Access

MIMIC-CXR-Ext-MIMIC-CXR-DB: Linking MIMIC-IV Chest Radiographs with Rich and Harmonized Clinical Context

Houcemeddine Turki ,  Lukman Ismaila ,  Ahmed Ben Salem ,  Ahmed Nebli ,  Mahdi Kchaou ,  Abdulhameed Dere ,  Anas Alzahrani ,  Makram Koubaa

Published Oct. 1, 2026 · Version 1.0
When using this resource, please cite:

Turki, H., Ismaila, L., Ben Salem, A., Nebli, A., Kchaou, M., Dere, A., Alzahrani, A., & Koubaa, M. (2026). MIMIC-CXR-Ext-MIMIC-CXR-DB: Linking MIMIC-IV Chest Radiographs with Rich and Harmonized Clinical Context (version 1.0). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/gy0s-e371

Please include the standard citation for PhysioNet:

Pollard, T., Moody, B. E., Lehman, L., Gow, B., Fernandes, C., Xie, C., Johnson, A., Mark, R. G., & Heldt, T. (2026). PhysioNet as a global platform for biomedical research. Nature Health. https://doi.org/10.1038/s44360-026-00096-z. Available from: https://rdcu.be/faatM

Abstract

MIMIC-CXR-DB is a relational clinical database comprising 146,333 chest radiographs linked to hospital admissions, designed to support research at the intersection of electronic health records (EHRs) and chest radiography. The database integrates hospital admissions data, demographic characteristics, diagnoses, laboratory measurements, radiology metadata, radiographic findings, and discharge documentation. It is derived primarily from MIMIC-IV and MIMIC-CXR-JPG, enabling admission-level and imaging-based multimodal analyses. Radiographic findings are derived from chest X-ray reports in the MIMIC-CXR-JPG dataset and include standardized indicators of common thoracic conditions such as atelectasis, cardiomegaly, consolidation, edema, pneumonia, pneumothorax, and pleural abnormalities, alongside detailed DICOM-derived imaging metadata. Laboratory data are normalized using dictionary tables to enable consistent interpretation of test results across admissions. Clinical findings from free-text radiology and discharge reports are extracted as triples using Named Entity Recognition and Ontology-Based Entity Linking and represented in two dedicated tables to improve the discoverability of clinical insights. All data are de-identified and derived from the MIMIC ecosystem under established privacy and data-access requirements. By providing admission-level linkage between clinical variables and imaging studies, MIMIC-CXR-DB enables reproducible research in clinical epidemiology, radiology, and machine learning, particularly for the development and evaluation of models that combine electronic health records with medical imaging.


Background

The past decade has seen an unprecedented growth in the availability and scale of medical imaging datasets, driven by advances in digital radiology infrastructure, machine learning, and the increasing demand for automated clinical decision support systems. Among clinical imaging modalities, chest radiography (chest X-ray) remains the most commonly performed exam worldwide due to its low cost, wide accessibility, and utility in diagnosing a spectrum of cardiopulmonary conditions, from pneumonia and pneumothorax to heart failure and lung nodules. Chest X-rays are, therefore, a central target for both clinical research and computational analysis [1].

Large-scale, publicly available datasets have become foundational to advances in medical image analysis. Historically, early efforts such as the NIH ChestX-ray14 dataset, comprising over 110,000 frontal chest X-rays annotated with 14 disease labels, enabled benchmark comparisons and spurred research on deep learning-based classification [2]. Subsequently, datasets like CheXpert (~224,000 images) expanded the scope of pathology labels and introduced uncertainty modeling, while PadChest and other international cohorts diversified imaging sources and diagnostic taxonomies [3]. Most recently, even larger initiatives such as ReXGradient-160K have emerged, offering paired radiographic images and free-text reports across tens of thousands of patients to support next-generation models for automated report generation and foundation model training [4]. Despite these expansions, challenges remain in label quality, generalization across populations, and robustness to domain shift—limitations that are actively studied in the literature and highlight the need for diverse, multimodal datasets with rich clinical context [5].

The MIMIC Chest X-ray (MIMIC-CXR) dataset represents a seminal contribution within this landscape. It contains hundreds of thousands of de-identified radiographic studies linked to free-text radiology reports, permitting a wide range of tasks from disease classification and report generation to fine-grained report mining and image interpretation under real-world clinical variation [1, 6]. However, purely imaging datasets often lack overt linkage to other facets of clinical care, such as laboratory results, admission context, and longitudinal outcomes. This gap constrains the ability of models to integrate imaging with structured clinical data, which is increasingly recognized as essential for clinically meaningful predictions and integrated decision support.

To address this need, MIMIC-CXR-DB extends the imaging resources of MIMIC-CXR by integrating them with structured clinical data (hospital admissions, lab measurements, diagnoses, and discharge reports), forming a comprehensive relational database that supports multimodal research. This integration situates radiographic findings within the broader clinical trajectory of patients, enabling not only image-based modeling but also joint analysis of imaging signatures alongside laboratory biomarkers, diagnosis codes, and patient outcomes. In doing so, MIMIC-CXR-DB aligns with the state of the art in medical data science, where multimodal datasets that combine imaging with rich clinical context are increasingly valued for advancing robust, clinically relevant machine learning models.


Methods

Main Data Extraction and Processing

Data were extracted from the MIMIC-IV 3.1 [7] and MIMIC-CXR-JPG [6] datasets hosted in Google BigQuery. Access was obtained through an authenticated Google Cloud session within Google Colab, and queries were executed using the BigQuery Python client. The primary query integrated CheXpert-labeled chest radiographs (mimic_cxr_jpg.chexpert) with corresponding DICOM metadata (mimic_cxr_jpg.metadata) via subject and study identifiers. Each radiograph was then linked to hospital admissions (mimiciv_3_1_hosp.admissions) when the study timestamp fell within the admission–discharge interval.

Demographic variables were incorporated from the patients table (mimiciv_3_1_hosp.patients), and ICD-coded diagnoses were appended using left joins to diagnoses_icd and d_icd_diagnoses (mimiciv_3_1_hosp). Redundant columns were removed after loading the joined dataset into pandas. For each admission, diagnosis codes were consolidated into a single primary diagnosis (sequence number = 1) and a list of secondary diagnoses (sequence number > 1). These structured diagnosis fields were merged back into the imaging dataset, duplicates were removed, and the final table, containing CheXpert labels, imaging metadata, admission information, demographics, and diagnoses, was exported as an Excel file for downstream analyses.

Laboratory Examinations

Candidate laboratory tests relevant to chest radiograph interpretation were identified by screening the laboratory dictionary (mimiciv_3_1_hosp.d_labitems) and selecting clinically meaningful items through physician review. The curated list of ITEMIDs (common_labitems.xlsx) was used to extract laboratory measurements from mimiciv_3_1_hosp.labevents via BigQuery. Laboratory events were linked to admissions and, subsequently, to imaging studies when the radiograph timestamp occurred within the corresponding admission period.

For each admission–lab item pair, the earliest recorded value was selected using a window function. Extracted values included the raw measurement, units, and reference ranges, and were categorized as “Higher,” “Normal,” or “Lower” according to documented reference limits. After removing intermediate processing columns, the finalized laboratory dataset, containing one representative value per admission per test, was exported to Excel.

Named Entity Recognition of Free-Text Notes

Free-text clinical notes were obtained by querying mimic_cxr_jpg.metadata, mimiciv_3_1_hosp.admissions, mimiciv_note.discharge, and mimiciv_note.radiology through an authenticated BigQuery session. Imaging studies were matched to discharge and radiology notes when the study timestamp fell within the corresponding admission interval. Retrieved tables were loaded into pandas DataFrames, exported to Excel, and reloaded for downstream text processing.

A custom parser converted each note into a hierarchical dictionary by identifying section headers formatted as all-capital lines or colon-delimited labels (e.g., IMPRESSION:, PHYSICAL EXAM:). Subsections were detected, and multiline content was aggregated under the appropriate sections. Parsed notes were transformed into a long format by exploding dictionary keys and values, enabling section-level frequency quantification. For discharge notes, only section types appearing at least 100 times during initial parsing and 500 times during NER preprocessing were retained; no frequency thresholds were applied to radiology reports.

Named entity recognition was performed using scispaCy with the en_core_sci_lg model augmented with the NegEx pipeline (negspacy) to identify negated expressions. All sections containing free text were processed, and for each entity we recorded its text span, negation status, subject identifier, hospital admission ID, and standardized section label. Entity extraction was parallelized using swifter, and the resulting entity tables were exported as pipe-delimited files.

Entity Linking

Ontology-based grounding of extracted entities was performed using a set of biomedical ontologies: SNOMED CT [8], OCHV [9], MeSH [10], PathLex [11], RadLex [12], SOHO [13], GALEN [14], LOINC [15], UPHENO [16], DCM [17], SYMP [18], and IOBC [19]. SNOMED CT was generated locally from RF2 release files using ROBOT and rdflib; all other ontologies were obtained from BioPortal [20].

Entities were resolved sequentially, beginning with a combined SNOMED CT + OCHV search and proceeding through the remaining ontologies until a match was found. Candidate concepts were identified using fuzzy string matching (RapidFuzz) [21] against all lexical forms of each ontology entry (e.g., rdfs:label, skos:prefLabel, skos:altLabel). A minimum similarity threshold of 90% was required for assignment. When no match met this threshold, the entity text was iteratively truncated by removing the first or last word, and the ontology search was repeated until a match was identified or the term could no longer be shortened.

The resulting table represents each entry in a subject–property–object format, where the subject is the hospital admission (for discharge notes) or the radiology study (for radiology reports), the property corresponds to the note section, and the object is the recognized concept. This work is grounded in the assumption that section titles encode contextual information that clarifies the semantic roles of named entities within the associated sections [22].

Dictionary Creation and Data Cleaning

We standardized heterogeneous properties, such as section titles, by mapping them to a curated ontology of clinically meaningful categories, for example grouping variants of “PHYSICAL EXAM,” “HPI,” or “PAST MEDICAL HISTORY” under unified parent labels. This standardization was performed by physicians with assistance from ChatGPT 4.0. For the resolution of recognized concepts, we created a dictionary table capturing the main labels (rdfs:label) and the first-order superclass (i.e., the direct subclass of owl:Thing or equivalent) for each concept. We use ChatGPT 4.0 to homogenize the classes in the dictionary table, ensuring they reflect the primary categories of medical concepts. To ensure quality, we retained only the best associations between properties and classes: 100 associations for discharge notes and 50 associations for radiology reports. All final associations were validated by physicians, who removed any unrelated or spurious mappings. Concepts removed during validation were subsequently deleted from the dictionary tables to maintain a high-quality, curated mapping.

Data Cleaning using TypeSafe's Jev

This stage applies TypeSafe’s Jev as an automated semantic and biomedical plausibility validation layer to the MIMIC-CXR-DB data, transforming the original open-vocabulary extraction into a cleaner proof-of-concept dataset. After merging and deduplicating relation–concept pairs, Jev first evaluates whether each relation is meaningful and biologically plausible, then checks whether the associated concept represents a valid biomedical entity. Only relation–label keys that pass both checks are retained, resulting in 3,268 validated radiology keys (71.62%) and 12,155 validated discharge keys (83.69%). This process substantially reduces noisy or weakly specific relations while preserving provenance and providing a reproducible, automated filtering step without replacing expert review or clinical validation.


Data Description

MIMIC-CXR-DB is a relational clinical database designed to support research at the intersection of electronic health records (EHRs) and chest radiography. It integrates hospital admissions data, laboratory measurements, radiology metadata, radiographic findings, and discharge documentation. The database is derived primarily from MIMIC-IV clinical data and MIMIC-CXR-JPG, enabling longitudinal, admission-level, and imaging-based analyses.

Key features include:

  • Admission-level demographic and diagnostic data
  • Structured laboratory test results
  • Radiographic findings extracted from chest X-ray reports
  • Metadata linking radiology studies to DICOM images and external resources

All identifiers are de-identified and compliant with HIPAA Safe Harbor requirements.

Cohort Size

For the MIMIC-IV v3.1-based construction allowing the creation of this database:

  • 377,095 radiographs were considered.
  • 227,827 imaging studies were represented.
  • 146,333 radiographs (38.8%) could be placed within an inpatient admission window.
  • Radiographs that could not be linked to an inpatient admission are retained, with a null hadm_id.

This distinction is important when designing analyses. The admission-linked subset should not automatically be interpreted as representative of the entire MIMIC-CXR population, since the linkage procedure selects studies occurring within inpatient admission windows.

Table Descriptions

1. admissions.csv

Purpose

Contains hospital admission–level demographic, administrative, and diagnostic information for patients who have associated chest radiographs.

Primary Key

hadm_id – Unique hospital admission identifier

Columns

Column Name Description
hadm_id Unique identifier for a hospital admission
admittime Date and time of hospital admission
dischtime Date and time of hospital discharge
deathtime Date and time of in-hospital death (NULL if patient survived)
admission_type Type of admission (e.g., EMERGENCY, ELECTIVE, URGENT)
admission_location Location from which the patient was admitted
discharge_location Location to which the patient was discharged
marital_status Marital status at admission
race Patient-reported race or ethnicity
gender Patient gender
anchor_age Age of patient at admission (de-identified anchor age)
Primary Diagnosis Primary diagnosis associated with the admission
Other Diagnoses Secondary diagnoses, separated by semicolons

Notes

  • One row per hospital admission
  • hadm_id links to labs, radiographs, and discharge

2. d_concepts.csv

Purpose

A dictionary table defining clinical or radiological concepts used elsewhere in the database.

Primary Key

id

Columns

Column Name Description
id Unique concept identifier
label Human-readable concept name
class Concept category or ontology class

Notes

  • Used for normalization and semantic consistency
  • May reference radiographic findings, diagnoses, or ontology mappings

3. d_labs.csv

Purpose

Defines laboratory test metadata, serving as a lookup table for laboratory measurements.

Primary Key

ITEMID

Columns

Column Name Description
ITEMID Unique laboratory test identifier
LABEL Name of the laboratory test
CATEGORY Category of the lab test (e.g., Chemistry, Hematology)
FLUID Specimen type (e.g., Blood, Urine, Plasma)

Notes

  • Links to labs.itemid
  • Derived from MIMIC laboratory dictionaries

4. labs.csv

Purpose

Stores quantitative laboratory results measured during a hospital admission.

Primary Key

Composite key (hadm_id, itemid, timestamp not provided)

Columns

Column Name Description
hadm_id Hospital admission identifier
itemid Laboratory test identifier (links to d_labs.ITEMID)
valuenum Numeric value of the lab test
valueuom Unit of measurement
Status Result status (e.g., Final, Corrected)

Notes

  • One admission may have multiple lab results
  • Temporal ordering may require external timestamps if available

5. radiographs.csv

Purpose

Contains radiographic study metadata and structured imaging findings derived from chest X-ray reports. Extracted from MIMIC-CXR-JPG v2.1.0 (Available on PhysioNet).

Primary Key

dicom_id (image-level) study_id (study-level)

Columns

Clinical Finding Labels (Binary Indicators)

Column Name Description
Atelectasis Presence of atelectasis
Cardiomegaly Presence of cardiomegaly
Consolidation Presence of lung consolidation
Edema Presence of pulmonary edema
Enlarged_Cardiomediastinum Enlarged cardiomediastinal silhouette
Fracture Bone fracture
Lung_Lesion Lung lesion
Lung_Opacity Non-specific lung opacity
No_Finding No abnormal findings detected
Pleural_Effusion Pleural effusion
Pleural_Other Other pleural abnormality
Pneumonia Pneumonia
Pneumothorax Pneumothorax
Support_Devices Presence of medical support devices

Imaging and Metadata Fields

Column Name Description
subject_id Unique patient identifier
study_id Radiology study identifier
dicom_id Unique image identifier
study_datetime Date and time of imaging study
StudyDate DICOM study date
StudyTime DICOM study time
PerformedProcedureStepDescription Description of performed imaging procedure
ViewPosition X-ray view (e.g., PA, AP, LATERAL)
Rows Image height in pixels
Columns Image width in pixels
ProcedureCodeSequence_CodeMeaning DICOM procedure meaning
ViewCodeSequence_CodeMeaning DICOM view meaning
PatientOrientationCodeSequence_CodeMeaning Patient orientation
hadm_id Associated hospital admission

Notes

  • Labels are derived from NLP-extracted radiology reports
  • Values for radiological signs typically encoded as "1.0" for positive mentions, "0.0" for negative mentions, "-1.0" for unsure mentions, and "null" for no mentions
  • One admission may link to multiple studies and images

6. radiology.csv (or radiology_jev.csv for Jev-cleaned edition)

Purpose

Maps radiology studies to clinical concepts, such as human diseases or anatomy structures.

Columns

Column Name Description
study_id Radiology study identifier
relation_type Type of clinical relationship (e.g., PAST MEDICAL HISTORY)
object_uri URI pointing to the clinical concept

Notes

  • Enables the modelling of clinical interpretations of medical images in the form of triples (study_id, relation_type, object_uri)

7. discharge.csv (or discharge_jev.csv for Jev-cleaned edition)

Purpose

Links hospital admissions (discharge reports) to clinical concepts, such as human diseases or anatomy structures.

Columns

Column Name Description
hadm_id Hospital admission identifier
relation_type Type of clinical relationship (e.g., Imaging and Diagnostics.General)
object_uri URI pointing to the clinical concept

Notes

  • Enables the modelling of clinical information in discharge reports in the form of triples (hadm_id, relation_type, object_uri)

Entity Relationships Summary

The Entity Relationship Diagram of the MIMIC-CXR-DB Database is available in ERD.png.


Usage Notes

Use of the dataset is free to all researchers after signing of a data use agreement which stipulates, among other items, that (1) the user will not share the data, (2) the user will make no attempt to reidentify individuals, and (3) any publication which makes use of the data will also make the relevant code available.

Intended Use Cases

  • Clinical outcome prediction using imaging + labs
  • Radiology label validation and benchmarking
  • Multimodal machine learning (EHR + imaging)
  • Epidemiological studies of chest pathology
  • Longitudinal patient trajectory analysis

Data Provenance and Compliance

  • Derived from MIMIC-IV and MIMIC-CXR-JPG
  • Fully de-identified
  • Approved for use under PhysioNet credentialed access
  • Not suitable for patient re-identification or clinical decision-making

Limitations

The human validation stage only validated the shape of the dataset (i.e., its schema, columns, classes, and data model). No human validation has been done to verify single claims beyond the Jev-based automated plausibility screening, which cannot replace a proper expert revision of the database (mainly radiology.csv, discharge.csv, and their Jev-based variants). Additionally, due to the non-usage of hadm_id in MIMIC-CXR-JPG, we failed to fully match chest x-rays to corresponding patient admissions and radiology reports, respectively in MIMIC-IV and MIMIC-IV-NOTE.


Release Notes

This is the initial release of MIMIC-CXR-DB, which undergoes high-level human validation. Several problems can exist in the data extracted from radiology and discharge free-text reports. A future direction of this work is to have the two generated tables proofread by medical experts, ensuring they can be reliably reused in medical image analysis.


Ethics

The dataset is a derivative dataset of MIMIC-IV and MIMIC-CXR-JPG and thus no new patient data was collected. The ethics approval of the dataset follows from that of the parent MIMIC dataset.


Acknowledgements

We thank Aiman Ghrab (University of Sfax, Tunisia), Ahmed Dhia El Euch (University of Sfax, Tunisia), Nassir Navab (Technical University of Munich, Germany), and Chantal Pellegrini (Technical University of Munich, Germany) for useful comments and discussion.


Conflicts of Interest

The authors have no conflicts of interest to declare.


References

  1. Johnson, A. E., Pollard, T. J., Berkowitz, S. J., Greenbaum, N. R., Lungren, M. P., Deng, C. Y., et al. (2019). MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data, 6(1), 317.
  2. Pang, T., Li, P., & Zhao, L. (2023). A survey on automatic generation of medical imaging reports based on deep learning. BioMedical Engineering OnLine, 22(1), 48.
  3. Merkow, J., Soin, A., Long, J., Cohen, J. P., Saligrama, S., Bridge, C., et al. (2023, October). CheXstray: a real-time multi-modal monitoring workflow for medical imaging AI. In International Conference on Medical Image Computing and Computer-Assisted Intervention (pp. 326-336). Cham: Springer Nature Switzerland.
  4. Zhang, X., Acosta, J. N., Miller, J., Huang, O., & Rajpurkar, P. (2025). ReXGradient-160K: A Large-Scale Publicly Available Dataset of Chest Radiographs with Free-text Reports. arXiv preprint arXiv:2505.00228.
  5. Rafferty, A., Ramaesh, R., & Rajan, A. (2025). Limitations of Public Chest Radiography Datasets for Artificial Intelligence: Label Quality, Domain Shift, Bias and Evaluation Challenges. arXiv preprint arXiv:2509.15107.
  6. Johnson, A. E., Pollard, T. J., Greenbaum, N. R., Lungren, M. P., Deng, C. Y., Peng, Y., et al. (2019). MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs. arXiv preprint arXiv:1901.07042.
  7. Johnson, A. E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., et al. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific data, 10(1), 1.
  8. Chang, E., & Mostafa, J. (2021). The use of SNOMED CT, 2013-2020: a literature review. Journal of the American Medical Informatics Association, 28(9), 2017-2026.
  9. Amith, M., Cui, L., Roberts, K., Xu, H., & Tao, C. (2019, November). Ontology of consumer health vocabulary: providing a formal and interoperable semantic resource for linking lay language and medical terminology. In 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) (pp. 1177-1178). IEEE.
  10. Lipscomb, C. E. (2000). Medical subject headings (MeSH). Bulletin of the Medical Library Association, 88(3), 265.
  11. Govender, L., Geitner, J., Tyam, N., Botha, F. C. J., & Yeats, J. (2021). Pathology Lexicon AZ: A multilingual glossary app. African Journal of Health Professions Education, 13(4), 212-213.
  12. Langlotz, C. P. (2006). RadLex: a new method for indexing online educational materials. Radiographics, 26(6), 1595-1597.
  13. Kollapally, N. M., Chen, Y., Xu, J., & Geller, J. (2022, December). An ontology for the social determinants of health domain. In 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) (pp. 2403-2410). IEEE.
  14. Jovic, A., Prcela, M., & Gamberger, D. (2007, June). Ontologies in medical knowledge representation. In 2007 29th International conference on information technology interfaces (pp. 535-540). IEEE.
  15. Bodenreider, O., Cornet, R., & Vreeman, D. J. (2018). Recent developments in clinical terminologies—SNOMED CT, LOINC, and RxNorm. Yearbook of medical informatics, 27(01), 129-139.
  16. Matentzoglu, N., Bello, S. M., Stefancsik, R., Alghamdi, S. M., Anagnostopoulos, A. V., Balhoff, J. P., et al. (2025). The Unified Phenotype Ontology: a framework for cross-species integrative phenomics. Genetics, 229(3), iyaf027.
  17. Clunie, D. A. (2021). DICOM format and protocol standardization—a core requirement for digital pathology success. Toxicologic Pathology, 49(4), 738-749.
  18. Mohammed, O., Benlamri, R., & Fong, S. (2012, December). Building a diseases symptoms ontology for medical diagnosis: an integrative approach. In The First International Conference on Future Generation Communication Technologies (pp. 104-108). IEEE.
  19. Kushida, T., Kozaki, K., Kawamura, T., Tateisi, Y., Yamamoto, Y., & Takagi, T. (2019). Interconnection of biological knowledge using NikkajiRDF and interlinking ontology for biological concepts. New Generation Computing, 37(4), 525-549.
  20. Turki, H., Pasha, N. A., Altammami, A., & Alzahrani, A. H. (2025). Automating Epidemiology Report Generation from the MIMIC-IV Clinical Database using SNOMED CT and SQL. Research Square.
  21. Ye, A., Wang, L., Zhao, L., Ke, J., Wang, W., & Liu, Q. (2021). RapidFuzz: Accelerating fuzzing via generative adversarial networks. Neurocomputing, 460, 195-204.
  22. Turki, H., Hadj Taieb, M. A., & Ben Aouicha, M. (2018). MeSH qualifiers, publication types and relation occurrence frequency are also useful for a better sentence-level extraction of biomedical relations. Journal of Biomedical Informatics, 83, 217-218.

Files