Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

369

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

369 results for “Datasets, Benchmarking”

Learn how ShareScore rates datasets ↗
zenodo36/100

Benchmark datasets to study fairness in synthetic data generation

<p>The traveltime dataset is based on the Folktables project covering US census data. The target is a binary variable encoding whether or not the individual needs to travel more than 20 minutes for work; here, having a shorter travel time is the desirable outcome. &nbsp;We use a subset of data from the states of California, Florida, Maine, New York, Utah, and Wyoming states in 2018. Although the folktables dataset does not have any missing values, there are some values recorded as NaN due to the Bureau's data collection methodology. We remove the "esp" column, which encodes the employment status of parents, and has 99.55% missing values. We encode the missing values in the povpip, income to poverty ratio (0.85%), to -1 in accordance to the methodology in Ding et al.. See https://arxiv.org/pdf/2108.04884 for metadata.</p> <p>The cardio (a) dataset contains patient data recorded during medical examination, including 3 binary features supplied by the patient. The target class denotes the presence of cardiovascular disease. This dataset represents predictive tasks that allocate access to priority medical care for patients, and has been used for fairness evaluations in the domain.</p> <p>The credit dataset contains historical financial data of borrowers, including past non-serious delinquencies. Here, a serious delinquency is considered to be 90 days past due, and this is the target variable.</p> <p>The German Credit dataset (https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data) contains financial and personal information regarding loan-seeking applicants.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Benchmark Dataset for Structure Refinement Methods of Protein Complex Models

<p>This is the dataset used in our work entitled &quot; Benchmarking of Structure Refinement Methods for Protein Complex Model &quot; by Jacob Verburgt and Daisuke Kihara, which is under review.</p> <p>ZDOCK Derived Benchmark Dataset:</p> <p>The primary benchmark set used in was directly derived from the <a href="https://zlab.umassmed.edu/benchmark/">ZDOCK Benchmark set</a>. The ZDOCK set contains four structures per target: An unbound ligand, an unbound receptor, a bound ligand, and a bound receptor. The benchmark set is available in such a way where the coordinates of the bound subunits are oriented identical to their complex structure, and the unbound subunits are superimposed onto their respective bound subunits. Our dataset creates the optimially oriented &quot;unbound&quot; complexes by combining the superimposed and unbound subunits, along with removal of waters, ligands, and other non-protein atoms. These unbound complexes are saved in the dataset in the form &quot;XXXX_c_u.pdb&quot;, where XXXX is the PDB ID.</p> <p>From the complete ZDOCK Benchmark of 230 targets, 18 targets were removed due to containing multiple ligand chains, which is incompatible with the standard ligand to receptor model used within CAPRI. The ZDOCK PDB ID&acirc;&euro;&trade;s of these targets are 1AKJ&quot;, &quot;1BJ1&quot;, &quot;1DE4&quot;, &quot;1EER&quot;, &quot;1EXB&quot;, &quot;1EZU&quot;, &quot;1GP2&quot;, &quot;1I9R&quot;, &quot;1JMO&quot;, &quot;1K74&quot;, &quot;1N2C&quot;, &quot;1QFW&quot;, &quot;2HMI&quot;, &quot;3EO1&quot;, &quot;3HMX&quot;, &quot;4FQI&quot;, &quot;4GXU&quot;, and &quot;9QFW&quot;.</p> <p>There are an additional 8 targets where the superimpostion of the ligand and receptor structures onto the complex led to entanglement of the chains and were subsequently removed from the dataset. The ZDOCK PDB IDs for these targets are &quot;1BGX&quot;, &quot;1H1V&quot;, &quot;1IRA&quot;, &quot;1R8S&quot;, &quot;1Y64&quot;, &quot;2OT3&quot;, &quot;3AAD&quot;, &quot;4GAM&quot;.</p> <p>Note:</p> <p>In the work, we also used CAPRI scoring model dataset derived from CAPRI rounds 38-45. This dataset is unable to be distributed directly by us due to CAPRI guidelines, but can be derived from &quot;Scoring round&quot; models from the <a href="https://www.ebi.ac.uk/pdbe/complex-pred/capri/">CAPRI Website </a>.</p> <p>The targets considered were T122-T125, T131-T133, and T136, as these were targets which contained globular protein ligands and receptors. Please contact us directly if you have any further questions on this dataset.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Benchmarking datasets for RiboDetector paper

<p>The benchmarking datasets used in RiboDetector paper, including five datasets described in the material and methods part of the paper</p>

opencc-by-sa-4.0Mar 2021View details →
dryad36/100

Mental health network of Gipuzkoa dataset (2015) for benchmark analysis

<p>This dataset contains data from the Mental Health System of Gipuzkoa, which are described in the manuscript: "Garcia-Alonso, CR., Almeda, N., Salinas-Pérez, JA., Gutiérrez-Colosía, MR., Iruin-Sanz, A., &amp; Salvador-Carulla., L. (2021). Use of a decision support system for benchmarking analysis and organizational improvement of regional mental health care: The case of Gipuzkoa (Basque Country, Spain)". This manuscript has been submitted to Plos One journal.</p> <p>This research focused on developing an analytical process for assessing the performance of the Mental health (MH) system of Gipuzkoa and identifying benchmark and target-for-improvement catchment areas. For doing so a decision support system that integrated data envelopment analysis, Monte Carlo simulation and artificial intelligence was used. The units of analysis, which are considered the decision-making units, were the 13 catchment areas defined by a reference MH centre. The following indicators were assessed: relative technical efficiency, stability and entropy to guide organizational interventions.</p> <p>Main results of the analyses pointed out that the MH system of Gipuzkoa showed high efficiency scores in each main type of care (inpatient, day and outpatient), but it can be considered unstable (small changes can have relevant impacts on MH provision and performance). With regards to performance improvement, it is recommended to reduce admissions and readmissions for inpatient care, increase workforce capacity and utilization of day care services and increase the availability of outpatient care services.</p>

opencc-zeroOct 2021View details →
zenodo36/100

CottonWeedDet12: a 12-class weed dataset of cotton production systems for benchmarking AI models for weed detection

<p>The dataset&nbsp;<strong>CottonWeedDet12</strong>&nbsp;consists of 5648 RGB images of 12-class&nbsp;weeds that are common in cotton fields in the southern U.S. states, with a total of 9370 bounding boxes. These images were acquired by either smartphones or hand-held digital cameras, under natural field light condition and throughout June to September of 2021. The images were manually labeled by qualified personnel for weed identification, and the labeling process was done using the VGG Image Annotator (version 2.10).</p> <p>The dataset, at the time of publication, is the largest publicly available multi-class dataset dedicated to weed detection. It expects to facilitate communicate efforts to exploit state-of-the-art deep learning method to push weed recognition to the next level. With the WeedDet12 dataset, a performance benchmark of a suite of YOLO object detectors has been built for weed detection. Detailed documentation of the dataset, model benchmarking and performance results is given in an accompanying journal paper: <a href="https://www.sciencedirect.com/science/article/pii/S0168169923000431">Dang, F., Chen, D., Lu, Y., Li, Z., 2023. YOLOWeeds: A novel benchmark of YOLO object detectors for multi-class weed detection in cotton production systems. Computers and Electronics in Agriculture 205, 107655. https://doi.org/10.1016/j.compag.2023.107655</a><a href="https://doi.org/10.1016/j.compag.2023.107655">&nbsp;</a></p> <p>If you use the dataset on a published publication, please cite the dataset or the <a href="https://doi.org/10.1016/j.compag.2023.107655">journal article</a> above.</p>

opencc-by-nc-4.0Jan 2023View details →
zenodo36/100

PolarBearVidID: A Video-based Re-Identification Benchmark Dataset for Polar Bears

<p><em><strong>The peer-reviewed publication for this dataset has now been published&nbsp;in&nbsp;Animals,&nbsp;an&nbsp;MDPI journal, and can be accessed here:&nbsp;<a href="https://doi.org/10.3390/ani13050801">https://doi.org/10.3390/ani13050801</a>. Please cite this when using the dataset.</strong></em></p> <p><em>PolarBearVidID</em>&nbsp;includes 13 individual polar bears housed in six institutions. Each identity has 110 sequences on average. The maximum length of the sequences is 8 seconds, respectively 100 frames at a frame rate of 12.5 frames per second. The average length of the sequences is 96.69 images. In total, the dataset includes 1431 sequences. The resolution of the images is set to 256 x 128 pixels. Finally, <em>PolarBearVidID</em>&nbsp;is the first dataset to enable utilizing the movement of a non-human species as a feature for the task of re-identification.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

The International FluidFlower benchmark study dataset

<p>The dataset describes physical, laboratory&nbsp;CO<sub>2</sub> injection experiments in a room-scale physcial reservoir model with geological realistic geometry. The dynamics of relevant subsurface CO2 injection and trapping mechanisms are recorded in time-lapsed image-series. Five repetitions of operationally identical CO2 injection experiments were performed. For each of the five repetitions (termed C1, C2, C3, C4 and C5) one dataset is issued. Each dataset contains 137 high-resolution images with the following intervals: 10 images before CO<sub>2</sub> injection at 20 second intervals; images every 5 min during the first 360 min (6 hours) of the experiment (73 images); images every hour until 48 hours (42 images); images every 6 hours until end of experiment (12 images).</p> <p>&nbsp;</p> <p>The format of image names is&nbsp;yyMMdd_timeHHmmss_image number.TIF (e.g. &#39;211124_time082740_DSC00067.TIF&#39;)</p> <p>Note: image series C5 contains 133 images; the 4 missing images from the 5-minute interval between 5-6 hours.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification

<p>EuroSAT is a land use and land cover classification dataset. The dataset is based on Sentinel-2 satellite imagery covering 13 spectral bands and consists&nbsp;of 10 LULC classes with a total of&nbsp;27,000 labeled and geo-referenced images. The dataset is associated with the publications &quot;<a href="https://ieeexplore.ieee.org/abstract/document/8519248">Introducing EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification</a>&quot; and &quot;<a href="https://ieeexplore.ieee.org/abstract/document/8736785">EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification</a>&quot;.</p> <p>EuroSAT_RGB.zip contains the RGB version of the dataset, which includes&nbsp;the optical R, G and B frequency bands encoded as JPEG images.</p> <p>EuroSAT_MS.zip contains the multi-spectral version of the EuroSAT dataset, which includes&nbsp;all 13 Sentinel-2 bands&nbsp;in the original value range.</p>

openmit-licenseJul 2018View details →
zenodo36/100

GROBID end-to-end benchmarking datasets

<p>Here are the datasets used for <a href="https://github.com/kermitt2/grobid">GROBID</a> end-to-end benchmarking covering:</p> <p>- metadata extraction,</p> <p>- bibliographical reference extraction, parsing and citation context identification, and</p> <p>- full text body structuring.</p> <p>The following collections are included:</p> <p>- a PubMedCentral gold-standard dataset called <code><strong>PMC_sample_1943</strong>, </code>compiled by Alexandru Constantin. The dataset, around 1.5GB in size, contains 1943 articles from 1943 different journals corresponding to the publications from a 2011 snapshot. For each article, we have a PDF file and a NLM XML file.</p> <p>- a bioRxiv dataset called&nbsp;<strong>biorxiv-10k-test-2000</strong> of 2000 preprint articles originally compiled with care and published by Daniel Ecer, available on <a href="https://zenodo.org/record/3873702">Zenodo</a>. The dataset contains for each article a PDF file and the corresponding reference NLM file (manually created by bioRxiv). The NLM files have been further systematically reviewed and annotated with additional markup corresponding to data and code availability statements and funding statements by the Grobid team. Around 5.4G in size.</p> <p>- a set of 1000 PLOS articles, called <strong>PLOS_1000,</strong> randomly selected from the full <a href="https://allof.plos.org/allofplos.zip">PLOS Open Access collection</a>. Again, for each article, the published PDF is available with the corresponding publisher JATS XML file, around 1.3GB total size.</p> <p>- a set of 984 articles from eLife, called <strong>eLife_984</strong>, randomly selected from their <a href="https://github.com/elifesciences/elife-article-xml">open collection</a> available on GitHub. Every articles come with the published PDF, the publisher JATS XML file and the eLife public HTML file (as bonus, not used), all in their latest version, around 4.5G total.</p> <p>For each of these datasets, the directory structure is the same and documented <a href="https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/#directory-structure">here</a>.</p> <p>Further information on Grobid benchmarking and how to run it: <a href="https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/">https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/</a>. Latest benchmarking scores are also available in the Grobid documentation: <a href="https://grobid.readthedocs.io/en/latest/Benchmarking/">https://grobid.readthedocs.io/en/latest/Benchmarking/</a></p> <p>These resources are originally published under CC-BY license. Our additional annotations are similarly under CC-BY.</p> <p>We thank NIH, bioRxiv, PLOS and eLife for making these resources Open Access and reusable.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Datasets for benchmarking RNA 2D structure prediction algorithms.

<p>Datasets for benchmarking ML approaches in RNA 2D structure prediction task.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

TauBench: A Dynamic Benchmark for Graphics Rendering (Dataset)

<p>TauBench is a dynamic graphics rendering benchmark dataset, targeted especially towards&nbsp;rendering methods relying on the reuse of temporal data. Path traced reference frames for the dataset are available at&nbsp;<a href="https://zenodo.org/record/7907164">https://zenodo.org/record/7907164</a>.</p> <p>The related conference paper (describing TauBench 1.0) is available at <a href="https://zenodo.org/record/6223036">https://zenodo.org/record/6223036</a>. More information about the 1.1 update and about the dataset in general can be found at&nbsp;<a href="https://webpages.tuni.fi/vga/taubench">https://webpages.tuni.fi/vga/taubench</a>.</p>

openother-ncMay 2023View details →
zenodo36/100

A Benchmark Dataset for Vision-based Bridge Traffic Load Monitoring in a Cable-stayed Bridge

<p>Traffic load monitoring based on deep learning and computer vision has garnered significant attention in bridge engineering worldwide. Unlike traditional traffic load monitoring systems, computer vision-based techniques can accurately extract the spatiotemporal load distribution across the entire bridge in an autonomous manner. However, many of the related studies in the literature used datasets that were collected from a few specific areas of different bridges, and there are very limited datasets that provide complete coverage of the entire bridge, making a detailed comparison of different methods difficult. This paper presents a benchmark dataset that provides a series of annotations and field measurements required for traffic load detection, tracking and continuous monitoring on the bridge. The dataset was collected by five cameras, and two weigh-in-motion systems installed on a cable-stayed bridge and is divided into three subsets. The first subset contains over 32,000 images and annotation files of eleven types of vehicle-related targets, which are necessary for the training of vehicle detection models. The second subset consists of photos of the calibration board and coordinates of reference points that are used for camera calibration. The last subset is designated for the field verification of various algorithms, providing synchronized vehicle weight data and monitoring videos covering the whole bridge. To the author&rsquo;s knowledge, this dataset is the first open-source dataset for vision-based traffic load monitoring in a bridge, which will have tremendous value in promoting research in the area of innovative bridge health monitoring technologies.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

BubbleML: A Multi-Physics Dataset and Benchmarks for Machine Learning

<p>In the field of phase change phenomena, the lack of accessible and diverse datasets suitable for machine learning (ML) training poses a significant challenge. Existing experimental datasets are often restricted, with limited availability and sparse ground truth data, impeding our understanding of this complex multi-physics phenomena. To bridge this gap, we present the <a href="https://github.com/HPCForge/BubbleML">BubbleML</a> Dataset &nbsp;which leverages physics-driven simulations to provide accurate ground truth information for various boiling scenarios, encompassing nucleate pool boiling, flow boiling, and sub-cooled boiling. This extensive dataset covers a wide range of parameters, including varying gravity conditions, flow rates, sub-cooling levels, and wall superheat, comprising 79 simulations. &nbsp;BubbleML is validated against experimental observations and trends, establishing it as an invaluable resource for ML research. Furthermore, we showcase its potential to facilitate exploration of diverse downstream tasks by introducing two benchmarks: (a) optical flow analysis to capture bubble dynamics, and (b) operator networks for learning temperature dynamics. The BubbleML dataset and its benchmarks serve as a catalyst for advancements in ML-driven research on multi-physics phase change phenomena, enabling the development and comparison of state-of-the-art techniques and models.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Dataset for Posiform Planting: Generating QUBO Instances for Benchmarking

<p>Dataset for the paper titled Posiform Planting: Generating QUBO Instances for Benchmarking</p> <p>https://arxiv.org/abs/2308.05859</p> <p>LA-UR-23-29274</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

choderalab/geometry-benchmark-espaloma: Small molecule geometry benchmark dataset to validate espaloma-0.3

<p>This is a collection of preprocessed QM and MM optimized structures needed to perform the small molecule geometry benchmark study, described in the <strong>espaloma-0.3</strong> paper:</p> <p>Kenichiro Takaba, Iv&aacute;n Pulido,&nbsp;Pavan Kumar Behara, Mike Henry, Hugo MacDermott Opeskin, John D. Chodera, Yuanqing Wang.&nbsp;&quot;Machine-learned molecular mechanics force field for the simulation of protein-ligand systems and beyond&quot; (<a href="https://arxiv.org/abs/2307.07085">arXiv:2307.07085</a>)</p> <p>This benchmark study calculates and compares the RMSD, TFD, and ddE metrics for a specified set of MM force fields. The initial optimized structures were&nbsp;sourced from the&nbsp;<a href="https://github.com/openforcefield/qca-dataset-submission/tree/master/submissions/2021-06-04-OpenFF-Industry-Benchmark-Season-1-v1.1">OpenFF Industry Benchmark Season 1 v1.1</a>&nbsp;dataset, which is available through &nbsp;<a href="https://qcarchive.molssi.org/">QCArchive</a>.&nbsp;More details about the preprocessing steps is available at&nbsp;<a href="https://github.com/choderalab/geometry-benchmark-espaloma/tree/main/qc-opt-geo">https://github.com/choderalab/geometry-benchmark-espaloma/tree/main/qc-opt-geo</a>.</p> <ul> <li><strong>02-chunks.tar.gz</strong>:&nbsp;QM optimized structures chunked into small file sizes.</li> <li><strong>02-outputs-openff-2.0.0-espaloma-0.3.0rc1.tar.gz</strong>:&nbsp;MM optimized structures using openff-2.0.0 and espaloma-0.3.0rc1 force field (former release candidate of espaloma-0.3)</li> <li><strong>02-outputs-gaff2.11.tar.gz</strong>:&nbsp;MM optimized structures using gaff-2.11 force field</li> <li><strong>02-outputs-espaloma-0.3.0rc6.tar.gz</strong>:&nbsp;MM optimized structures using espaloma-0.3.0rc6 (espaloma-0.3) force field</li> <li><strong>02-outputs-openff-2.1.0.tar.gz</strong>: MM optimized structures using openff-2.1.0 force field</li> </ul>

opencc-by-4.0Sep 2023View details →
zenodo36/100

UAV-PDD2023: A benchmark dataset for pavement distress detection based on UAV images

<p>The images in the dataset ( VOC format) were captured by a UAV at an altitude of 30 meters. The collected images were annotated in PASCAL VOC format. A total of 11,158&nbsp;instances in 2,440&nbsp;images are incorporated in the dataset.</p> <ul> <li>The UAV-PDD2023 dataset, captured by unmanned aerial vehicles (UAVs), provides a benchmark for road damage detection. It is highly useful for municipal authorities and road agencies to conduct low-cost road condition monitoring.&nbsp;</li> <li>Six types of road damages are labeled in the dataset: Longitudinal cracks (LC), Transverse cracks (TC), Alligator cracks (AC), Oblique cracks (OC), Repair (RP), and Potholes (PH).&nbsp;</li> <li>Researchers can use this dataset as a benchmark to evaluate the performance of different algorithms in addressing similar problems, such as image classification and object detection.&nbsp;</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

maDLC Marmoset Benchmark Dataset - Test

<p><a href="https://zenodo.org/record/5849371">see&nbsp;https://benchmark.deeplabcut.org/ for more information.</a></p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

maDLC Parenting Benchmark Dataset - Test set

<p>see&nbsp;https://benchmark.deeplabcut.org/ for more information.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

maDLC Fish Benchmark Dataset - Test set

<p><a href="https://zenodo.org/record/5849286">see benchmark.deeplabcut.org for more information.</a></p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Datasets used for benchmarking of JOINTLY.

<p>Datasets used for benchmarking of JOINTLY with associated labels, intermediary results and final results. Version 2 holds the datasets used in the final manuscript.</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record