Resources

47 results for "benchmark dataset"

Database Credentialed Access

MIMIC-Ext-MIMIC-CXR-VQA: A Complex, Diverse, And Large-Scale Visual Question Answering Dataset for Chest X-ray Images

Seongsu Bae, Daeun Kyung, Jaehee Ryu, et al.

We introduce MIMIC-Ext-MIMIC-CXR-VQA (i.e., Extended from MIMIC database), a complex, diverse, and large-scale dataset designed for Visual Question Answering (VQA) tasks within the medical domain, focusing primarily on chest radiographs. This datase…
Published: July 19, 2024. Version: 1.0.0
Database Credentialed Access

RuMedNLI: A Russian Natural Language Inference Dataset For The Clinical Domain

Pavel Blinov, Aleksandr Nesterov, Galina Zubkova, et al.

There is a shortage of text medical resources for the Russian language. This is a substantial obstacle in state-of-the-art NLP deep learning models research and development. To mitigate this issue we translated the MedNLI data from English to Russia…
Published: April 1, 2022. Version: 1.0.0
Database Credentialed Access

CardioWave-Portable: a portable ECG and PPG dataset for cuffless blood pressure estimation

Hailin Yang, Lexi Zhang, Hongyu Liu, et al.

CardioWave-Portable is a controlled-access dataset of synchronized single-lead electrocardiography (ECG) and photoplethysmography (PPG) signals acquired with a portable card-type device and paired with cuff-based systolic and diastolic blood pressur…
Published: Sept. 30, 2026. Version: 1.0.0
Database Open Access

NInFEA: Non-Invasive Multimodal Foetal ECG-Doppler Dataset for Antenatal Cardiology Research

Danilo Pani, Eleonora Sulas, Monica Urru, et al.

The development of algorithms for the extraction of the foetal ECG (fECG) from non-invasive recordings is hampered by the lack of publicly-available reference datasets, which could be used to benchmark different algorithms while providing a ground t…
Published: Nov. 12, 2020. Version: 1.0.0 | Visualize waveforms
Model Credentialed Access

Me-LLaMA: Foundation Large Language Models for Medical Applications

Qianqian Xie, Qingyu Chen, Aokun Chen, et al.

Recent advancements in large language models (LLMs) such as ChatGPT and LLaMA have hinted at their potential to revolutionize medical applications, yet their application in clinical settings often reveals limitations due to a lack of specialized tra…
Published: June 5, 2024. Version: 1.0.0
Database Restricted Access

Swiss-Mammo: A physician-written, synthetic dataset of German mammography reports

Daniel Reichenpfader, Sandro von Däniken, Harald Marcel Bonel

This dataset, Swiss-Mammo, contains 28 manually constructed German mammography reports, each paired with an English translation. The reports are stratified across BI-RADS categories 0 through 6, with four reports per category. All reports were manua…
Published: June 24, 2025. Version: 1.0.1
Database Credentialed Access

EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images

Seongsu Bae, Daeun Kyung, Jaehee Ryu, et al.

Electronic Health Records (EHRs), which contain patients' medical histories in various multi-modal formats, often overlook the potential for joint reasoning across imaging and table modalities underexplored in current EHR Question Answering (QA) sys…
Published: July 23, 2024. Version: 1.0.0
Database Credentialed Access

CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays

Hyungyung Lee, Geon Choi, Jung Oh Lee, et al.

Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks such as report generation and visual question answering. However, existing benchmarks focus mainly on the final diagnostic answer, offering l…
Published: Oct. 23, 2025. Version: 1.0.1
Database Open Access

Hillel Yaffe Glaucoma Dataset (HYGD): A Gold-Standard Annotated Fundus Dataset for Glaucoma Detection

Or Abramovich, Hadas Pizem, Jonathan Fhima, et al.

Glaucomatous optic neuropathy (GON) is a leading cause of irreversible blindness worldwide, affecting an estimated 64.3 million people globally with projections reaching 111.8 million by 2040. Approximately 50% of cases remain undiagnosed until adva…
Published: March 16, 2026. Version: 1.1.0
Database Credentialed Access

RadNLI: A natural language inference dataset for the radiology domain

Yasuhide Miura, Yuhao Zhang, Emily Tsai, et al.

The problem of natural language inference (NLI) determines whether a natural language hypothesis can be justifiably inferred from a natural language premise. NLI has attracted researchers to benchmark it in a number of settings including medical one…
Published: June 29, 2021. Version: 1.0.0