Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.7.1
Dataset results
11 results for “benchmark training dataset”
TocoDecoy: a new approach to design unbiased datasets for training and benchmarking machine-learning scoring functions
<p>This dataset file contains TocoDecoy datasets generated based on the targets and active ligands of LIT-PCBA.</p> <p>1_property_filtered.zip :</p> <ul> <li>TD set: the ligand file name, 2D T-sne vectors, Smiles, molecular weight (MW), Wildman-Crippen partition coefficient (log P), number of rotatable bonds (RB), number of hydrogen-bond acceptors (HBA), number of hydrogen-bond donors (HBD), number of halogens (HAL), topology similarities of decoys to the seed active ligands, active label (active or inactive) and training set label (whether belongs to training set or test set) <strong>OF active ligands and their topologically dissimilar decoys</strong></li> <li>CD set: the decoy conformations with low docking scores generated by docking active ligands into protein pockets using Glide, Schrödinger.</li> </ul> <p> </p>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Training Set Part 1)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>Due to the size limit of zenodo.org, we split the MELA training set into 3 parts; this is the Training Set Part 1 of MELA dataset, including 260 CTs. Files include:</p> <ol> <li>Train1.zip: 130 CTs in NII format (nii.gz).</li> <li>Train2.zip: 130 CTs in NII format (nii.gz).</li> </ol>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Training Set Part 2)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>Due to the size limit of zenodo.org, we split the MELA training set into 3 parts; this is the Training Set Part 2 of MELA dataset, including 260 CTs. Files include:</p> <ol> <li>Train3.zip: 130 CTs in NII format (nii.gz).</li> <li>Train4.zip: 130 CTs in NII format (nii.gz).</li> </ol>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Training Set Part 3)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>Due to the size limit of zenodo.org, we split the MELA training set into 3 parts; this is the Training Set Part 3 of MELA dataset, including 250 CTs. Files include:</p> <ol> <li>Train5.zip: 130 CTs in NII format (nii.gz).</li> <li>Train6.zip: 120 CTs in NII format (nii.gz).</li> </ol>
maDLC Tri-Mouse Benchmark Dataset - Training
<p>see https://benchmark.deeplabcut.org/ for more information.</p>
maDLC Marmoset Benchmark Dataset - Training
<p>see https://benchmark.deeplabcut.org/ for more information.</p>
maDLC Fish Benchmark Dataset - Training
<p>see benchmark.deeplabcut.org for more information.</p>
RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Training Set Part 2)
<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification. We hope this large-scale dataset could facilitate both clinical research for automatic rib fracture detection and diagnoses, and engineering research for 3D detection, segmentation and classification.</p> <p>Due to size limit of zenodo.org, we split the whole RibFrac Training Set into 2 parts; This is the Training Set Part 2 of RibFrac dataset, including 120 CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-train-images-2.zip: 120 chest-abdomen CTs in NII format (nii.gz).</li> <li>ribfrac-train-labels-2.zip: 120 annotations in NII format (nii.gz).</li> <li>ribfrac-train-info-2.csv: labels in the annotation NIIs. <ul> <li>public_id: anonymous patient ID to match images and annotations.</li> <li>label_id: discrete label value in the NII annotations.</li> <li>label_code: 0, 1, 2, 3, 4, -1 <ul> <li>0: it is background</li> <li>1: it is a displaced rib fracture</li> <li>2: it is a non-displaced rib fracture</li> <li>3: it is a buckle rib fracture</li> <li>4: it is a segmental rib fracture</li> <li>-1: it is a rib fracture, but we could not define its type due to ambiguity, diagnosis difficulty, etc. Ignore it in the classification task. </li> </ul> </li> </ul> </li> </ol> <p> </p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using bibtex</p> <p><em>@article{ribfrac2020,<br> title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> journal={EBioMedicine},<br> year={2020},<br> publisher={Elsevier}<br> }</em></p> <p> </p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p> </p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>
RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Training Set Part 1)
<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification. We hope this large-scale dataset could facilitate both clinical research for automatic rib fracture detection and diagnoses, and engineering research for 3D detection, segmentation and classification.</p> <p>Due to size limit of zenodo.org, we split the whole RibFrac Training Set into 2 parts; This is the Training Set Part 1 of RibFrac dataset, including 300 CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-train-images-1.zip: 300 chest-abdomen CTs in NII format (nii.gz).</li> <li>ribfrac-train-labels-1.zip: 300 annotations in NII format (nii.gz).</li> <li>ribfrac-train-info-1.csv: labels in the annotation NIIs. <ul> <li>public_id: anonymous patient ID to match images and annotations.</li> <li>label_id: discrete label value in the NII annotations.</li> <li>label_code: 0, 1, 2, 3, 4, -1 <ul> <li>0: it is background</li> <li>1: it is a displaced rib fracture</li> <li>2: it is a non-displaced rib fracture</li> <li>3: it is a buckle rib fracture</li> <li>4: it is a segmental rib fracture</li> <li>-1: it is a rib fracture, but we could not define its type due to ambiguity, diagnosis difficulty, etc. Ignore it in the classification task. </li> </ul> </li> </ul> </li> </ol> <p> </p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using bibtex</p> <p><em>@article{ribfrac2020,<br> title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> journal={EBioMedicine},<br> year={2020},<br> publisher={Elsevier}<br> }</em></p> <p> </p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p> </p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>
CodeEval: Pedagogy Based Benchmark Dataset for Evaluation Of Large Language Models Trained On Code
<p>Categories has been modified to 21 from 26.</p>
maDLC Parenting Benchmark Dataset - Training
<p>see https://benchmark.deeplabcut.org/ for more information.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.