Skip to main content
zenodoopen

A living catalogue of artificial intelligence datasets and benchmarks for medical decision making

<p>We provide&nbsp;a comprehensive curated catalogue of&nbsp;<strong>artificial intelligence datasets</strong> and <strong>benchmarks for medical decision making</strong>. At the time of first release (April 2021), the dataset contains more than 400&nbsp;biomedical and clinical datasets&nbsp;of which 252 are publicly available or available upon request.</p> <p>The dataset was compiled based on a systematic literature review covering both biomedical and computer science literature and&nbsp;grey literature data sources. All datasets were manually systematized and annotated for meta-information, such as:</p> <ul> <li>Availability and licensing information</li> <li>Type of source data</li> <li>Links to source publications, main references or dataset repositories</li> </ul> <p>Benchmark dataset were additionally annotated for the following information:</p> <ul> <li>Associated task</li> <li>Performance metrics commonly used for evaluation</li> <li>Clinical relevance</li> <li>The availability of data splits</li> </ul> <p>In addition to the versioned TSV file on Zenodo, the dataset can also be explored live via&nbsp;<a href="https://docs.google.com/spreadsheets/d/1QjUxxnZ3tuyW5dj6nkt_o5yJcWUZec4ttfJxO8Zlty4/edit?usp=sharing">this Google Spreadsheet</a>.&nbsp;The dataset is intended as a living, extendable resource. Edit suggestions and additions are encouraged and can be submitted via the comment function of the Google sheet.</p> <p>&nbsp;</p> <p><strong>File descriptions</strong></p> <p><em>annotated-datasets.tsv</em> -- contains the annotated datasets</p> <p><em>arXiv-literature-export.tsv</em> -- contains the original literature record export from arXiv</p> <p><em>pubmed-literature-export.tsv</em> -- contains the original literature record export from PubMed</p> <p><em>README.md</em> -- contains a detailed description of all annotation fields</p>

ShareScore

48/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
8

Topics