Resources

47 results for "benchmark dataset"

Database Credentialed Access

FFA-IR: Towards an Explainable and Reliable Medical Report Generation Benchmark

Mingjie Li, Wenjia Cai, Rui Liu, et al.

Automatic medical report generation (MRG) towards describing life-threatening lesions from given medical images, such as Chest X-ray and Fundus Fluorescein Angiography (FFA), has been a long-standing research topic in machine learning and automatic …
Published: Jan. 21, 2025. Version: 1.1.0
Database Credentialed Access

RadGraph2: Tracking Findings Over Time in Radiology Reports

Adam Dejl, Sameer Khanna, Patricia Therese Pile, et al.

RadGraph2 is a dataset of 800 chest radiology reports annotated using a fine-grained entity-relationship schema, which is an expanded version of the previously introduced RadGraph dataset. In contrast with the previous approaches and the original Ra…
Published: Aug. 8, 2024. Version: 1.0.0
Database Restricted Access

Pulmonary Edema Severity Grades Based on MIMIC-CXR

Ruizhi Liao, Geeticka Chauhan, Polina Golland, et al.

Clinical management decisions for patients with acutely decompensated heart failure and many other diseases are often based on grades of pulmonary edema severity, rather than its mere absence or presence. Chest radiographs are commonly performed to …
Published: Feb. 9, 2021. Version: 1.0.1
Database Open Access

A multimodal gait dataset of brain activity, muscle activity, kinematics and ground forces in young adults

Rateb Katmah, Aamna AlShehhi, Doua Kosaji, et al.

Gait is a fundamental motor function, and its analysis is essential for understanding locomotor control, rehabilitation, and the early detection of neurological and musculoskeletal disorders. While many datasets capture either biomechanical or neura…
Published: April 30, 2026. Version: 1.0.0 | Visualize waveforms
Database Credentialed Access

Medical-CXR-VQA dataset: A Large-Scale LLM-Enhanced Medical Dataset for Visual Question Answering on Chest X-Ray Images

Xinyue Hu, Lin Gu, Kazuma Kobayashi, et al.

Medical Visual Question Answering (VQA) is an important task in medical multi-modal Large Language Models (LLMs), aiming to answer clinically relevant questions regarding input medical images. This technique has the potential to improve the efficien…
Published: Jan. 21, 2025. Version: 1.0.0
Database Credentialed Access

MIMIC-III-Ext-Notes

Darren Liu, Monique Bouvier, Delgersuren Bold, et al.

Unstructured clinical documentation, such as progress notes, contains rich contextual information critical for clinical decision-making but remains underutilized in computational research due to the limited availability of annotated datasets. The MI…
Published: Feb. 27, 2026. Version: 1.0.0
Database Credentialed Access

MIMIC-III-Ext-MIMIC-Patient: Structured Per-Patient JSON Records for Clinical Question Answering

Tianqi Shang, Weiqing He, Charles Zheng, et al.

This project provides MIMIC-Patient, a structured, per-patient representation of the MIMIC-III database designed for large language model (LLM)–based clinical question answering. For each of 500 admissions, we reconstruct the patient’s episode into …
Published: Sept. 11, 2026. Version: 1.0.0
Database Credentialed Access

MIMIC-IV-ECHO-Ext-LVVOLUMES-A4C-ROI: Annotated Subset of Apical Four-Chamber Echocardiography for PoCUS-Style LV Volume and Function Analysis

Kamlin Ekambaram, Anurag Arnab, Philip Herbst, et al.

MIMIC-IV-ECHO-Ext-LVVOLUMES-A4C-ROIis a curated, credentialed subset of MIMIC-IV-ECHO v0.1 focusing on apical four-chamber (A4C) echocardiographic views with accompanying left-ventricular volumetric labels. The release includes 1,064 de-identified A…
Published: Feb. 26, 2026. Version: 1.0.0
Database Restricted Access

EchoNext: A Dataset for Detecting Echocardiogram-Confirmed Structural Heart Disease from ECGs

Pierre Elias, Joshua Finer

This dataset contains a de-identified collection of 100,000 12-lead electrocardiograms (ECGs) with paired structural heart disease (SHD) labels derived from echocardiography, collected at Columbia University Irving Medical Center. Each ECG is provid…
Published: April 30, 2026. Version: 1.1.1
Database Credentialed Access

MIMIC-IV-Ext-MedicalBench: Evaluating Large Language Models Towards Improved Medical Concept Extraction

Zhichao Yang, Gregory Lyng, Sanjit Batra, et al.

Medical concept extraction from electronic health records underpins many downstream applications, yet remains challenging because medically meaningful concepts, such as diagnosis, are frequently implied rather than explicitly stated in medical narra…
Published: March 23, 2026. Version: 1.0.0