Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

56

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

56 results for “Benchmark study”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data set supplementing "Benchmarking triage capability of symptom checkers against that of medical laypersons: Survey study"

<p>This is the de-identified data set used to conduct the analyses in the study published as Original Research in the JMIR under the title &quot;Benchmarking triage capability of symptom checkers against that of medical laypersons: Survey study&quot; (https://doi.org/10.2196/24475)</p> <p>The data set contains the assessments of the urgency of symptoms to 45 fictitious clinical case vignettes&nbsp;by&nbsp;91 US participants, and the participants&#39; age, gender and level of education. Data for the symptom checker apps is needed to fully reproduce our study and can be found in the appendix of the paper &quot;Evaluation of symptom checkers for self diagnosis and triage: audit study&quot;&nbsp;by Semigran et al. (2015) (https://doi.org/10.1136/bmj.h3480).</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena subject to imposed damage

<p>A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena (HBTA), a full-scale steel bridge subject to imposed damage, has been established. The data set includes organized dynamic response and load measurement data of the bridge under different structural state conditions, where the structural state conditions range from an undamaged (reference) state to known damage states. Furthermore, the data set includes acceleration and strain data from the response monitoring and acceleration data from the load monitoring, where a modal vibration shaker is used as an excitation source. The data is collected in one h5-file (hierarchical data format version 5) with a sampling rate of 100 Hz. Signal processing and resampling of the data has been performed according to the description provided in the references below. The data set is now published in this open-access data repository and can be accessed and downloaded freely. As such, the data set provides an important benchmark to the scientific community within bridge damage detection and SHM.</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Datasets used in the benchmarking study of MR methods

<p>We conducted a benchmarking analysis of 16 summary-level data-based MR methods for causal inference with five real-world genetic datasets, focusing on three key aspects: type I error control, the accuracy of causal effect estimates, replicability, and power.</p> <p>The datasets used in the MR benchmarking study can be downloaded here:</p> <ol> <li>"dataset-GWASATLAS-negativecontrol.zip":&nbsp; the GWASATLAS dataset for evaluation of type I error control in confounding scenario (a): Population stratification</li> <li>"dataset-NealeLab-negativecontrol.zip": the Neale Lab dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-PanUKBB-negativecontrol.zip": the Pan UKBB dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-Pleiotropy-negativecontrol": the dataset&nbsp; used for evaluation of type I error control in confounding scenario (b): Pleiotropy;</li> <li>"dataset-familylevelconf-negativecontrol.zip": the dataset used for evaluation of type I error control in confounding scenario (c): Family-level confounders;</li> <li>"dataset_ukb-ukb.zip": the dataset used for evaluation of the accuracy of causal effect estimates;</li> <li>"dataset-LDL-CAD_clumped.zip": the dataset used for evaluation of replicability and power;</li> </ol> <p>Each of the datasets contains the following files:</p> <ol> <li>&nbsp;"Tested Trait pairs": the exposure-outcome trait pairs to be analyzed;</li> <li>"MRdat" refers to the summary statistics after performing IV selection (p-value &lt; 5e-05) and PLINK LD clumping with a clumping window size of 1000kb and an r^2 threshold of 0.001.</li> <li>"bg_paras" are the estimated background parameters "Omega" and "C" which will be used for MR estimation in MR-APSS.</li> </ol> <p>Note:</p> <ol> <li>The formatted dataset after quality control can be accessible at our GitHub website (https://github.com/YangLabHKUST/MRbenchmarking).</li> <li>The details on quality control of GWAS summary statistics, formatting GWASs, and LD clumping for IV selection can be found on the MR-APSS software tutorial on the MR-APSS&nbsp;&nbsp;website (https://github.com/YangLabHKUST/MR-APSS).</li> <li>R code for running MR methods is also available at https://github.com/YangLabHKUST/MRbenchmarking.</li> </ol>

opencc-by-4.0Jan 2024View details →
zenodo44/100

SPEChpc 2021 Benchmarks: A Performance and Energy Case Study

SPEChpc 2021 Benchmarks on Ice Lake and Sapphire Rapids based Infiniband Clusters: A Performance and Energy Case Study.

opengpl-2.0Aug 2023View details →
zenodo40/100

Data Instances for: Who moves the locker? A benchmark study of alternative mobile parcel locker concepts

<p>|C|_h.txt</p> <p>|C|:&nbsp;&nbsp; &nbsp;number of customers<br> h: &nbsp;&nbsp; &nbsp;instance</p> <p>|C|;|P|;</p> <p>|C|:&nbsp;&nbsp; &nbsp;number of customers<br> |P|:&nbsp;&nbsp; &nbsp;number of parking spaces</p> <p>Customer (c;size;max_dist;min_time;L;x_1;y_1;...;x_L;y_L;s_1;e_1;...;s_L;e_L)</p> <p>c:&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;customer index&nbsp;<br> size: &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;parcel size<br> max_dist:&nbsp;&nbsp; &nbsp;maximum walking distance<br> min_time:&nbsp;&nbsp; &nbsp;minimum overlap time<br> L:&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;number of whereabouts<br> (x_i,y_i):&nbsp;&nbsp; &nbsp;position of whereabouts i<br> [s_i,e_i]:&nbsp;&nbsp; &nbsp;time window of whereabouts i</p> <p>Parking space (p;x;y)<br> p:&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;parking space index&nbsp;<br> (x,y):&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;position</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Datasets for "Precision and accuracy of single-molecule FRET measurements – a multi-laboratory benchmark study"

<p>Supplementary material (raw data) for Fig. 2 in &quot;<strong>Precision and accuracy of single-molecule FRET measurements &ndash; a multi-laboratory benchmark study</strong>&quot; to be published with Nature Methods</p> <p>The confocal data is given in ht3 and hdf5 format.</p> <p>For the TIRF data the original TIFF-stacks are uploaded including the calibration files.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

A boreal forest model benchmarking dataset for North America: a case study with the Canadian Land Surface Scheme including Biogeochemical Cycles (CLASSIC)

<p>A boreal forest model benchmarking dataset for North America by harmonizing eddy covariance and supporting measurements from black spruce (Picea mariana)-dominated mature forest stands.</p> <p>Dataset glossary and users&rsquo; instructions are documented in &lsquo;README.md&rsquo;.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Virtual sensors for wind energy applications benchmark study data - preliminary version

<p>Test version of the time series data for the wind energy virtual sensing benchmark study data.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Bugs4Q: A Benchmark of Existing Bugs to Enable Controlled Testing and Debugging Studies for Quantum Programs

<p>Realistic benchmarks of reproducible bugs and fixes are vital to good experimental evaluation of debugging and testing approaches.&nbsp;Bugs4Q is a benchmark of forty-two real, manually validated Qiskit bugs from three popular platforms (GitHub, StackOverflow, and Stack Exchange) in programming, supplemented with test cases to reproduce buggy behaviors.&nbsp;Bugs4Q Database allows users to access the bugs we collected directly. Bugs4Q Framework provides interfaces for accessing the buggy and fixed versions of the Qiskit programs and executing the corresponding source code and unit tests, facilitating reproducible empirical studies and comparisons of Qiskit program debugging and testing tools.</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Artifact Description/Artifact Evaluation/Computational Artifact for "SPEChpc 2021 Benchmarks on Ice Lake and Sapphire Rapids Infiniband Clusters: A Performance and Energy Case Study"

<p>We provide reproducibility initiative dependencies (Artifact Description or Artifact Evaluation or Computational Results Analysis) appendix at https://github.com/RRZE-HPC/PMBS23-AD. To allow a third party to duplicate the findings, this article provides our extensive performance data artifact and describes further details regarding the software environments, experimental design, and methodology employed for the results shown in the paper, entitled &quot;SPEChpc 2021 Benchmarks on Ice Lake and Sapphire Rapids Infiniband Clusters: A Performance and Energy Case Study&quot;. The computational artifacts will enable experienced performance engineers to reproduce and interpret the data shown in the paper in the appropriate way and to follow the conclusions we draw from it.</p>

opengpl-2.0Aug 2023View details →
zenodo40/100

Virtual sensors benchmark study test timeseries tier 0 - part 1

<p>A dataset with time series for testing virtual sensor models for wind turbine aeroelastic loads. Data are in zipped format, with 100 time series available.</p> <p>The test sets in the study are organized with several levels of &quot;difficulty&quot; according to added uncertainties and noise with respect to the training data. This particular dataset (tier 0) is with exactly the same distribution as the training data.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Virtual benchmarking study training time series - binary set 1

<p>A dataset with time series for training virtual sensor models for wind turbine aeroelastic loads. Data are in zipped format, each zip file contains 1000 individual time series. For the purpose of space preservation, each time series is stored in parquet binary format with &quot;snappy&quot; encoding.</p> <p>Reading a parquet file in Python can be done with the pandas library, with the following command sequence:</p> <p>import pandas as pd<br> Data = pd.read_parquet(filename)</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data for the paper "Insights gained from a comprehensive all-against-all transcription factor binding motif benchmarking study".

<p>Data&nbsp;for the&nbsp;paper &quot;Insights gained from a comprehensive all-against-all transcription factor binding motif benchmarking study&quot;.</p>

openmit-licenseMar 2020View details →
zenodo36/100

Experimental Data Sets for the study "Benchmarking a $(\mu+\lambda)$ Genetic Algorithm with Configurable Crossover Probability"

<p>This is the experimental result of the study &quot;Benchmarking a (&mu;+&lambda;) Genetic Algorithm with Configurable Crossover Probability&quot;. A novel&nbsp;(&mu;+&lambda;) GA is proposed and benchmarked, in which we stochastically determine whether to apply the crossover operator either for each individual or generation with a crossover probability&nbsp;<span class="math-tex">\(p_c\)</span>.&nbsp;This data set&nbsp;consists&nbsp;of two parts:</p> <ol> <li>The results of (&mu;+&lambda;) GA on 25 pseudo-Boolean problems defined in <em>IOHprofiler </em>(<a href="https://iohprofiler.github.io/">https://iohprofiler.github.io/</a>) with the following&nbsp;setup:&nbsp;<span class="math-tex">\(\mu \in \{10, 50, 100\}, \lambda \in \{1, \lceil\mu/2\rceil, \mu\}, p_c\in\{0, 0.5\}.\)</span> <ul> <li>&#39;IOHprofiler_Problems_standard_bit_mutation.csv&#39; --&gt; the (&mu;+&lambda;) GA with standard bit mutation.</li> <li>&#39;IOHprofiler_Problems_fast_mutation.csv&#39; --&gt; the (&mu;+&lambda;) GA with fast&nbsp;mutation.</li> </ul> </li> <li>The results of (&mu;+&lambda;) GA on OneMax and LeadingOnes problems&nbsp;with the following setup:&nbsp;<span class="math-tex">\(n \in \{64,100,150,200,250,500\}, \mu \in \{2,3,5,8,10,20,30,...,100\}, \\ \lambda \in \{1, \lceil \mu/2 \rceil, \mu\}, \text{and }p_c \in \{0.1 k \mid k \in [0..9]\}\cup\{0.95\}.\)</span> <ul> <li>&#39;OneMax_raw.csv&#39; --&gt; the fixed-target running time/first hitting time from 100 independent runs for target values in&nbsp;<span class="math-tex">\([1..n]\)</span>.</li> <li>&#39;OneMax_summary.csv&#39; --&gt; the mean, median, standard deviation, some quantiles, expected running time (ERT), the number of successful runs, and the success rate&nbsp;from 100 independent runs for target values in&nbsp;<span class="math-tex">\([1..n]\)</span>.</li> <li>&#39;LeadingOnes_raw.csv&#39; --&gt; the same with &#39;OneMax_raw.csv&#39; for LeadingOnes.</li> <li>&#39;LeadingOnes_summary.csv&#39; --&gt; the same with &#39;OneMax_summary.csv&#39; for LeadingOnes.</li> </ul> </li> </ol> <p><strong>Contact</strong>: if you have any questions or suggestions, please feel free to contact&nbsp;<a href="https://www.universiteitleiden.nl/en/staffmembers/furong-ye#tab-1">Furong Ye</a> or <a href="http://www-ia.lip6.fr/~doerr/">Carola Doerr</a>.</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Performance Measurement Dataset of the HPC Benchmarks FASTEST, Kripke, LULESH, MiniFE, Quicksilver, and RELeARN for Scalability Studies with Extra-P

<p>This dataset contains performance measurements of the HPC benchmarks FASTEST, Kripke, LULESH, MiniFE, Quicksilver, and RELeARN intended to be used for scalability studies with Extra-P (https://github.com/extra-p/extrap). The datasets contains measurements of various application configurations considering several model parameters, e.g., the number of MPI ranks and the input problem size, using weak scaling for each benchmark.</p>

openbsd-3-clauseNov 2023View details →
zenodo36/100

A benchmark study of ab initio gene prediction methods in diverse eukaryotic organisms

<p>G3PO (Gene and Protein Prediction PrOgrams) Benchmark was designed to represent many of the typical challenges faced by current genome annotation projects. The benchmark is based on a carefully validated and curated set of real eukaryotic genes from 147 phylogenetically disperse organisms (from human to protists).&nbsp;<br> &nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Benchmarking Study of Deep Generative Models for Inverse Polymer Design: Reinforcement Learning

<p>Well-trained models and generation results for reinforcement learning part of <a href="https://github.com/ytl0410/Polymer-Generative-Models-Benchmark">ytl0410/Polymer-Generative-Models-Benchmark: Well-trained models and generative outcomes for the paper "Benchmarking Study of Deep Generative Models for Inverse Polymer Design" (github.com)</a></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Estudios de referencia que se basaron en el DUA y cómo se empleó este enfoque / Benchmark studies that were based on the UDL and how this approach was used

<p>Una d&eacute;cada de investigaci&oacute;n sobre la eficacia de las pr&aacute;cticas de educaci&oacute;n inclusiva y el DUA en la universidad. Revisi&oacute;n sistem&aacute;tica de la literatura.<br>A decade of research on the effectiveness of inclusive education practices and UDL in universities. Systematic literature review.</p> <p><br>Mar&iacute;a Pineda-Mart&iacute;nez<br>Agosto de 2024</p> <p><br>Tabla/Table<br>Estudios de referencia que se basaron en el DUA y c&oacute;mo se emple&oacute; este enfoque / Benchmark studies that were based on the UDL and how this approach was used</p> <p><br>Principios, pautas y puntos de verificaci&oacute;n del DUA (versi&oacute;n 2.2.) identificados en la revisi&oacute;n sistem&aacute;tica / Principles, guidelines and checkpoints of the SAD&nbsp;(version 2.2.) identified in the systematic review.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data for "A realistic benchmark for differential abundance testing and confounder adjustment in human microbiome studies"

<p>Data for the manuscript: A realistic benchmark for differential abundance testing and confounder adjustment in human microbiome studies&nbsp;(see also https://doi.org/10.1101/2022.05.09.491139)</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Benchmark datasets to study fairness in synthetic data generation

<p>The traveltime dataset is based on the Folktables project covering US census data. The target is a binary variable encoding whether or not the individual needs to travel more than 20 minutes for work; here, having a shorter travel time is the desirable outcome. &nbsp;We use a subset of data from the states of California, Florida, Maine, New York, Utah, and Wyoming states in 2018. Although the folktables dataset does not have any missing values, there are some values recorded as NaN due to the Bureau's data collection methodology. We remove the "esp" column, which encodes the employment status of parents, and has 99.55% missing values. We encode the missing values in the povpip, income to poverty ratio (0.85%), to -1 in accordance to the methodology in Ding et al.. See https://arxiv.org/pdf/2108.04884 for metadata.</p> <p>The cardio (a) dataset contains patient data recorded during medical examination, including 3 binary features supplied by the patient. The target class denotes the presence of cardiovascular disease. This dataset represents predictive tasks that allocate access to priority medical care for patients, and has been used for fairness evaluations in the domain.</p> <p>The credit dataset contains historical financial data of borrowers, including past non-serious delinquencies. Here, a serious delinquency is considered to be 90 days past due, and this is the target variable.</p> <p>The German Credit dataset (https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data) contains financial and personal information regarding loan-seeking applicants.</p>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record