PhysioNet Index

Database Credentialed Access

Northwestern ICU (NWICU) database

Dana Moukheiber, William Temps, Bhadrappa Molgi, et al.

A freely available COVID-rich ICU database comprising de-identified health-related data from Northwestern Memorial Health Center (NHMC).

Published: Nov. 19, 2024. Version: 0.1.0

Model Credentialed Access

Shareable Artificial Intelligence to Extract Cancer Outcomes from Electronic Health Records for Precision Oncology Research

Kenneth Kehl, Pavel Trukhanov, Christopher Fong, et al.

The DFCI-imaging-student and DFCI-medonc-student AI models for extracting cancer outcomes from imaging reports and medical oncologist notes from electronic health records.

Published: Oct. 24, 2024. Version: 1.0.0

Database Contributor Review

CARMEN-I: A resource of anonymized electronic health records in Spanish and Catalan for training and testing NLP tools

Eulalia Farre Maduell, Salvador Lima-Lopez, Santiago Andres Frid, et al.

CARMEN-I is a Spanish corpus of 2,000 clinical records from Hospital Clínic, Barcelona. It covers COVID-19 patients and comorbidities, serving as a resource for training clinical NLP models and researchers in NLP applied to clinical documents.

de-identification clinical ner anonymization

Published: April 20, 2024. Version: 1.0.1

Database Open Access

MIMIC-IV Clinical Database Demo

Alistair Johnson, Lucas Bulgarelli, Tom Pollard, et al.

An openly available subset of patients in the MIMIC-IV database.

critical care electronic health record mimic

Published: Jan. 31, 2023. Version: 2.2

Database Credentialed Access

Learning to Ask Like a Physician: a Discharge Summary Clinical Questions (DiSCQ) Dataset

Eric Lehman

Dataset of questions asked by medical experts about patients. Medical experts will read a discharge summary line-by-line and (1) ask any question that they may have and (2) record what in the text "triggered" them to ask their question.

question generation question answering machine learning

Published: July 28, 2022. Version: 1.0

Database Credentialed Access

Critical care database comprising patients with infection at Zigong Fourth People's Hospital

Ping Xu, Lin Chen, Zhongheng Zhang

Routinely collected data from critical care units at Zigong Fourth People’s Hospital, Sichuan, China for patients admitted between January 2019 and December 2020 Missing information on temperature are updated in the new version.

Published: June 30, 2022. Version: 1.1

Database Credentialed Access

RadNLI: A natural language inference dataset for the radiology domain

Yasuhide Miura, Yuhao Zhang, Emily Tsai, et al.

A radiology NLI dataset introduced in the paper: Improving Factual Completeness and Consistency of Image-to-text Radiology Report Generation

Published: June 29, 2021. Version: 1.0.0

Database Credentialed Access

Paediatric Intensive Care database

Haomin Li, Xian Zeng, Gang Yu

PIC (Paediatric Intensive Care) is a large paediatric-specific, single-centre, bilingual database comprising information relating to children admitted to critical care units at a large children’s hospital in China.

intensive care pediatrics critical care natural language processing

Published: Nov. 12, 2020. Version: 1.1.0

Database Restricted Access

Gout Emergency Department Chief Complaint Corpora

John David Osborne, Tobias O'Leary, Amy Mudano, et al.

A corpus of chief complaints tagged with predicted gout flare status and chart reviewed gout flare status. Ideal for input to masked language model training to supplement lengthy clinical text notes.

gout emergency department nlp

Published: Oct. 19, 2020. Version: 1.0

Database Credentialed Access

MIMIC-III - SequenceExamples for TensorFlow modeling

Jonas Kemp, Kun Zhang, Andrew Dai

MIMIC-III data converted into TensorFlow SequenceExample format, for use in modeling pipelines.

tensorflow sequence modeling deep learning machine learning

Published: Sept. 29, 2020. Version: 1.0.0

Search

Resources

Northwestern ICU (NWICU) database

Shareable Artificial Intelligence to Extract Cancer Outcomes from Electronic Health Records for Precision Oncology Research

CARMEN-I: A resource of anonymized electronic health records in Spanish and Catalan for training and testing NLP tools

MIMIC-IV Clinical Database Demo

Learning to Ask Like a Physician: a Discharge Summary Clinical Questions (DiSCQ) Dataset

Critical care database comprising patients with infection at Zigong Fourth People's Hospital

RadNLI: A natural language inference dataset for the radiology domain

Paediatric Intensive Care database

Gout Emergency Department Chief Complaint Corpora

MIMIC-III - SequenceExamples for TensorFlow modeling