Resources


Database Credentialed Access

Insulin4RL: Real-Time Insulin Infusions For Offline Reinforcement Learning

Thomas Frost, Steve Harris

Openly available research dataset intended for offline reinforcement learning (ORL) using natively irregular healthcare data. The dataset is intended to encourage further research into ORL methods using naturally sporadic decision intervals.

insulin intensive care semi-markov decision process diabetes blood glucose offline reinforcement learning machine learning

Published: June 15, 2026. Version: 1.0.0


Database Restricted Access

EchoNext: A Dataset for Detecting Echocardiogram-Confirmed Structural Heart Disease from ECGs

Pierre Elias, Joshua Finer

EchoNext is a curated dataset of electrocardiograms (ECGs) paired with echocardiogram-confirmed structural heart disease labels, designed to support the development and validation of machine learning models.

clinical decision support artificial intelligence digital health structural heart disease electrocardiogram health equity ecg heart failure transthoracic echocardiogram ai model deployment valvular heart disease cardiovascular screening ai in healthcare left ventricular dysfunction deep learning population health aortic stenosis machine learning

Published: April 30, 2026. Version: 1.1.1


Database Open Access

A multimodal gait dataset of brain activity, muscle activity, kinematics and ground forces in young adults

Rateb Katmah, Aamna AlShehhi, Doua Kosaji, et al.

Dataset of gait recordings from 59 healthy adults, combining brain activity, muscle activity, body kinematics, and ground forces during treadmill walking at three different speeds.

electroencephalography biomechanics neuroscience gait analysis force plate kinematics inertial measurement unit electromyography

Published: April 30, 2026. Version: 1.0.0

Visualize waveforms

Database Open Access

Hillel Yaffe Glaucoma Dataset (HYGD): A Gold-Standard Annotated Fundus Dataset for Glaucoma Detection

Or Abramovich, Hadas Pizem, Jonathan Fhima, et al.

HYGD is a rigorously annotated fundus image dataset with gold-standard clinical labels designed to improve and benchmark deep learning models for accurate glaucoma detection.

ophthalmology retina glaucoma dfi gon fundus gold-standard

Published: March 16, 2026. Version: 1.1.0


Database Credentialed Access

MIMIC-III-Ext-Notes

Darren Liu, Monique Bouvier, Delgersuren Bold, et al.

We evaluated general large language models' performance in clinical information extraction on MIMIC-III notes.

Published: Feb. 27, 2026. Version: 1.0.0


Database Credentialed Access

MIMIC-IV-ECHO-Ext-LVVOLUMES-A4C-ROI: Annotated Subset of Apical Four-Chamber Echocardiography for PoCUS-Style LV Volume and Function Analysis

Kamlin Ekambaram, Anurag Arnab, Philip Herbst, et al.

A curated subset of MIMIC-IV-ECHO providing apical four-chamber cine loops with manual ROI masks, volumetric labels, and ready-to-use MP4/NPZ derivatives for robust LV volume and ejection fraction research.

ultrasound echocardiography medical imaging dicom lvesv roi segmentation cardiac video analysis left ventricular volume mimic-iv-echo apical four-chamber quantitative cardiology biplane simpson transformer models lvef ejection fraction a4c pocus lvedv domain adaptation deep learning

Published: Feb. 26, 2026. Version: 1.0.0


Database Credentialed Access

RadGraph-XL: A Large-Scale Expert-Annotated Dataset for Entity and Relation Extraction from Radiology Reports

Jean-Benoit Delbrouck

RadGraph-XL is a large, expert-annotated dataset of 2,300 radiology reports covering multiple modalities and anatomies. It enables accurate extraction of clinical entities and relations for downstream medical AI tasks.

Published: Sept. 12, 2025. Version: 1.0.0


Database Open Access

MIMIC-IV Clinical Database Demo on FHIR

Alex Bennett, Hannes Ulrich, Joshua Wiedekopf, et al.

The MIMIC-IV Clinical Database Demo on FHIR is a 100 patient subset of the MIMIC-IV v2.2 and MIMIC-IV-ED v2.2 clinical databases converted into the Fast Healthcare Interoperability Resources (FHIR) format.

fhir electronic health records mimic

Published: Aug. 27, 2025. Version: 2.1.0


Database Restricted Access

Swiss-Mammo: A physician-written, synthetic dataset of German mammography reports

Daniel Reichenpfader, Sandro von Däniken, Harald Marcel Bonel

Swiss-Mammo: A physician-written, synthetic dataset of 28 German mammography reports. The dataset is stratified based on BI-RADS categories and available in German and English.

radiology mammography structured reporting bi-rads

Published: June 24, 2025. Version: 1.0.1


Database Credentialed Access

Medical-CXR-VQA dataset: A Large-Scale LLM-Enhanced Medical Dataset for Visual Question Answering on Chest X-Ray Images

Xinyue Hu, Lin Gu, Kazuma Kobayashi, et al.

Medical-CXR-VQA provides a large-scale LLM-enhanced dataset for visual question answering in medical chest x-ray images.

Published: Jan. 21, 2025. Version: 1.0.0