Resources
53 results for "benchmark"
Database
Credentialed Access
CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays
Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks such as report generation and visual question answering. However, existing benchmarks focus mainly on the final diagnostic answer, offering l…
Database
Contributor Review
ER-REASON: A Benchmark Dataset for LLM-Based Clinical Reasoning in the Emergency Room
The ER-Reason dataset is a benchmark designed to evaluate LLM-based clinical reasoning and decision-making in the emergency room (ER), a high-stakes setting where clinicians make rapid, consequential decisions across diverse patient presentations an…
Database
Credentialed Access
FFA-IR: Towards an Explainable and Reliable Medical Report Generation Benchmark
Automatic medical report generation (MRG) towards describing life-threatening lesions from given medical images, such as Chest X-ray and Fundus Fluorescein Angiography (FFA), has been a long-standing research topic in machine learning and automatic …
Database
Credentialed Access
MIMIC-IV-Ext-MDS-ED: Multimodal Decision Support in the Emergency Department - a Benchmark Dataset for Diagnoses and Deterioration Prediction in Emergency Medicine
Measurable progress in the development of medical decision support systems has been hindered by a lack of comprehensive datasets. Many available datasets focus on narrow prediction tasks and do not include a diverse range of data types, which limits…
Database
Credentialed Access
EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries
Discharge summaries in Electronic Health Records (EHRs) are crucial for clinical decision-making, but their length and complexity make information extraction challenging, especially when dealing with accumulated summaries across multiple patient adm…
Database
Credentialed Access
Lunguage: A Benchmark for Structured and Sequential Chest X-ray Interpretation
Radiology reports convey detailed clinical observations and capture diagnostic reasoning that evolves over time. However, existing evaluation methods are limited to single-report settings and rely on coarse metrics that fail to capture fine-grained …
fine-grained structured reports
attribute-level clinical reasoning
medical text structuring
longitudinal clinical reasoning
chest x-ray report parsing
medical information structuring
benchmark dataset for radiology report
medical information extraction
structured radiology reports
temporal relation extraction
radiology report benchmarking
longitudinal clinical understanding
Database
Credentialed Access
MedVAL-Bench: Expert-Annotated Medical Text Validation Benchmark
MedVAL-Bench is a dataset containing physician evaluations of errors in language model (LM)-generated medical text. The dataset spans 6 diverse medical text generation tasks and includes annotations from 12 physicians on clinically significant error…
Database
Credentialed Access
MIMIC-IV-ECHO-Ext-MIMICEchoQA: A Benchmark Dataset for Echocardiogram-Based Visual Question Answering
We present MIMICEchoQA, a benchmark dataset for echocardiogram-based question answering, built from the publicly available MIMIC-IV-ECHO database. Each echocardiographic study was paired with the closest discharge summary within a 7-day window, and …
Database
Credentialed Access
CXR-Align: A Benchmark for CXR-Report Alignment with Negations
CXR-Align is a benchmark dataset designed to evaluate vision-language processing (VLP) models' ability to accurately interpret negations in chest X-ray (CXR) reports. Negations are prevalent in medical documentation and pose significant challenges f…
Database
Credentialed Access
ODD: A Benchmark Dataset for the NLP-based Opioid Related Aberrant Behavior Detection
Opioid related aberrant behaviors (ORAB) present novel risk factors for opioid overdose. Previously, ORAB have been mainly assessed by survey results and by monitoring drug administrations. Such methods however, cannot scale up and do not cover the …