Resources

9 results for "multi-label"

Challenge Credentialed Access

CXR-LT: Multi-Label Long-Tailed Classification on Chest X-Rays

Gregory Holste, Mingquan Lin, Song Wang, et al.

Chest radiography presents a "long-tailed" distribution of findings, where a few diseases are common, but most are rare. Diagnosis is further complicated by its multi-label nature, as patients often exhibit multiple co-occurring findings. While rece…
Published: March 19, 2025. Version: 2.0.0
Database Credentialed Access

A Brazilian Multilabel Ophthalmological Dataset (BRSET)

Luis Filipe Nakayama, Mariana Goncalves, Lucas Zago Ribeiro, et al.

The Brazilian Multilabel Ophthalmological Dataset (BRSET) is a multi-labeled ophthalmological dataset designed to improve scientific community development and validate machine learning models. In ophthalmology, ancillary exams support medical decisi…
Published: July 27, 2026. Version: 1.0.2
Database Credentialed Access

RadVLM Instruction Dataset

Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruiperez-Campillo, et al.

We release the RadVLM instruction dataset, a large-scale resource used to train the RadVLM model on diverse radiology tasks. The dataset contains 1,115,021 image–instruction pairs spanning five task families: (i) report generation from frontal CXRs …
Published: Sept. 25, 2025. Version: 1.0.0
Database Open Access

A Multi-Night Instantaneous Heart Rate and Accelerometry Dataset with EEG Sleep Stage Labels

Tzu-An Song, Yubo Zhang, Ziyuan Zhou, et al.

We collected sleep data from 47 healthy adult volunteers with no history of sleep disorders, recruited from the local community through study advertisement flyers. The study protocol was approved by the University of Massachusetts Lowell Institution…
Published: Sept. 25, 2026. Version: 1.0.1
Database Credentialed Access

MS-CXR-T: Learning to Exploit Temporal Structure for Biomedical Vision-Language Processing

Shruthi Bannur, Stephanie Hyland, Qianchu Liu, et al.

MS-CXR-T is a multi-modal benchmark dataset for evaluating biomedical vision-language processing (VLP) models on two distinct temporal tasks in radiology: image classification and sentence similarity. The former comprises multi-image frontal chest X…
Published: March 17, 2023. Version: 1.0.0
Database Credentialed Access

MIMIC-IV-ECHO-Ext-MIMICEchoQA: A Benchmark Dataset for Echocardiogram-Based Visual Question Answering

Rahul Thapa, Andrew Li, Qingyang Wu, et al.

We present MIMICEchoQA, a benchmark dataset for echocardiogram-based question answering, built from the publicly available MIMIC-IV-ECHO database. Each echocardiographic study was paired with the closest discharge summary within a 7-day window, and …
Published: Oct. 7, 2025. Version: 1.0.0
Database Credentialed Access

RadGraph: Extracting Clinical Entities and Relations from Radiology Reports

Saahil Jain, Ashwin Agrawal, Adriel Saporta, et al.

RadGraph is a dataset of entities and relations in full-text radiology reports. We designed a novel information extraction (IE) schema to structure clinical information in a radiology report with four entities and three relations. Our train set cons…
Published: June 3, 2021. Version: 1.0.0
Database Credentialed Access

Medical-CXR-VQA dataset: A Large-Scale LLM-Enhanced Medical Dataset for Visual Question Answering on Chest X-Ray Images

Xinyue Hu, Lin Gu, Kazuma Kobayashi, et al.

Medical Visual Question Answering (VQA) is an important task in medical multi-modal Large Language Models (LLMs), aiming to answer clinically relevant questions regarding input medical images. This technique has the potential to improve the efficien…
Published: Jan. 21, 2025. Version: 1.0.0
Database Credentialed Access

MIMIC-Ext-DrugDetection

Fabrice Harel-Canada, Nanyun Peng, David Goodman, et al.

This project shares a large, annotated drug detection dataset created from MIMIC-III/IV discharge summaries. The dataset was developed to address the challenge of identifying substance use behaviors in Electronic Health Records (EHRs), where critica…
Published: Sept. 25, 2025. Version: 1.0.0