Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11 results for “benchmark training dataset”

Learn how ShareScore rates datasets ↗
zenodo40/100

TocoDecoy: a new approach to design unbiased datasets for training and benchmarking machine-learning scoring functions

<p>This dataset file contains TocoDecoy datasets generated based on the targets and active ligands of LIT-PCBA.</p> <p>1_property_filtered.zip :</p> <ul> <li>TD set: the ligand file name, 2D T-sne vectors, Smiles, molecular weight (MW), Wildman-Crippen partition coefficient (log P), number of rotatable bonds (RB), number of hydrogen-bond acceptors (HBA), number of hydrogen-bond donors (HBD), number of halogens (HAL), topology similarities of decoys to the seed active ligands, active label (active or inactive) and training set label (whether belongs to training set or test set) <strong>OF active ligands and their topologically dissimilar decoys</strong></li> <li>CD set: the decoy conformations with low docking scores generated by docking active ligands into protein pockets using Glide, Schr&ouml;dinger.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Training Set Part 1)

<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset&nbsp;could facilitate the research and application of automatic mediastinal lesion detection and diagnosis.&nbsp;</p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>Due to the size limit of zenodo.org, we split the MELA training set into 3 parts; this is the Training Set Part 1 of MELA dataset, including 260 CTs. Files include:</p> <ol> <li>Train1.zip: 130 CTs in NII format (nii.gz).</li> <li>Train2.zip: 130 CTs in NII format (nii.gz).</li> </ol>

opencc-by-4.0Apr 2022View details →
zenodo40/100

MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Training Set Part 2)

<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset&nbsp;could facilitate the research and application of automatic mediastinal lesion detection and diagnosis.&nbsp;</p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>Due to the size limit of zenodo.org, we split the MELA training set into 3 parts; this is the Training Set Part 2 of MELA dataset, including 260 CTs. Files include:</p> <ol> <li>Train3.zip: 130 CTs in NII format (nii.gz).</li> <li>Train4.zip: 130 CTs in NII format (nii.gz).</li> </ol>

opencc-by-4.0Apr 2022View details →
zenodo40/100

MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Training Set Part 3)

<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset&nbsp;could facilitate the research and application of automatic mediastinal lesion detection and diagnosis.&nbsp;</p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>Due to the size limit of zenodo.org, we split the MELA training set into 3 parts; this is the Training Set Part 3&nbsp;of MELA dataset, including 250 CTs. Files include:</p> <ol> <li>Train5.zip: 130 CTs in NII format (nii.gz).</li> <li>Train6.zip: 120 CTs in NII format (nii.gz).</li> </ol>

opencc-by-4.0Apr 2022View details →
zenodo32/100

maDLC Tri-Mouse Benchmark Dataset - Training

<p>see&nbsp;https://benchmark.deeplabcut.org/ for more information.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

maDLC Marmoset Benchmark Dataset - Training

<p>see&nbsp;https://benchmark.deeplabcut.org/ for more information.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

maDLC Fish Benchmark Dataset - Training

<p>see benchmark.deeplabcut.org for more information.</p>

opencc-by-4.0Jan 2022View details →
zenodo28/100

RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Training Set Part 2)

<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification.&nbsp;We hope this large-scale dataset could facilitate both clinical&nbsp;research for automatic rib fracture detection and diagnoses, and engineering research&nbsp;for 3D detection, segmentation and classification.</p> <p>Due to size limit of zenodo.org, we split the whole RibFrac Training Set into 2 parts; This is the Training Set Part 2 of RibFrac dataset, including 120 CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-train-images-2.zip: 120 chest-abdomen&nbsp;CTs in NII format (nii.gz).</li> <li>ribfrac-train-labels-2.zip: 120 annotations in NII format (nii.gz).</li> <li>ribfrac-train-info-2.csv: labels in the annotation NIIs. <ul> <li>public_id: anonymous patient ID to match images and annotations.</li> <li>label_id: discrete label value&nbsp;in the NII annotations.</li> <li>label_code: 0, 1, 2, 3, 4, -1 <ul> <li>0: it is background</li> <li>1: it is a displaced rib fracture</li> <li>2: it is a non-displaced rib fracture</li> <li>3: it is a buckle rib fracture</li> <li>4: it is a segmental rib fracture</li> <li>-1: it is a rib fracture,&nbsp; but we could not define its type due to ambiguity, diagnosis difficulty, etc. Ignore it in the classification task.&nbsp;</li> </ul> </li> </ul> </li> </ol> <p>&nbsp;</p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using&nbsp; bibtex</p> <p><em>@article{ribfrac2020,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;journal={EBioMedicine},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;year={2020},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;publisher={Elsevier}<br> }</em></p> <p>&nbsp;</p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p>&nbsp;</p> <p>This work is licensed under a&nbsp;<a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>

opencc-by-nc-4.0Jun 2020View details →
zenodo28/100

RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Training Set Part 1)

<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification.&nbsp;We hope this large-scale dataset could facilitate both clinical&nbsp;research for automatic rib fracture detection and diagnoses, and engineering research&nbsp;for 3D detection, segmentation and classification.</p> <p>Due to size limit of zenodo.org, we split the whole RibFrac Training Set into 2 parts; This is the Training Set Part 1&nbsp;of RibFrac dataset, including 300&nbsp;CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-train-images-1.zip: 300 chest-abdomen CTs in NII format (nii.gz).</li> <li>ribfrac-train-labels-1.zip: 300&nbsp;annotations in NII format (nii.gz).</li> <li>ribfrac-train-info-1.csv: labels in the annotation NIIs. <ul> <li>public_id: anonymous patient ID to match images and annotations.</li> <li>label_id: discrete label value&nbsp;in the NII annotations.</li> <li>label_code: 0, 1, 2, 3, 4, -1 <ul> <li>0: it is background</li> <li>1: it is a displaced rib fracture</li> <li>2: it is a non-displaced rib fracture</li> <li>3: it is a buckle rib fracture</li> <li>4: it is a segmental rib fracture</li> <li>-1: it is a rib fracture,&nbsp; but we could not define its type due to ambiguity, diagnosis difficulty, etc. Ignore it in the classification task.&nbsp;</li> </ul> </li> </ul> </li> </ol> <p>&nbsp;</p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using&nbsp; bibtex</p> <p><em>@article{ribfrac2020,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;journal={EBioMedicine},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;year={2020},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;publisher={Elsevier}<br> }</em></p> <p>&nbsp;</p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p>&nbsp;</p> <p>This work is licensed under a&nbsp;<a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>

opencc-by-nc-4.0Jun 2020View details →
zenodo28/100

CodeEval: Pedagogy Based Benchmark Dataset for Evaluation Of Large Language Models Trained On Code

<p>Categories has been modified to 21 from 26.</p>

opencc-by-4.0May 2024View details →
zenodo24/100

maDLC Parenting Benchmark Dataset - Training

<p>see&nbsp;https://benchmark.deeplabcut.org/ for more information.</p>

opencc-by-4.0Jan 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record