Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

85

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

85 results for “standard dataset”

Learn how ShareScore rates datasets ↗
zenodo40/100

Players' Profiles and Preferences in Standard and Adapted Levels: Dataset

<p>Dataset produced in a study to measure the impact of levels&#39; generation and adaptation to the players&#39; preferences.</p> <p>The study was conducted in the context of a Master&#39;s thesis on Game Adaptivity.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Hunting for vampires and other unlikely forms of parity violation at the Large Hadron Collider: calo-image datasets for the standard model

<p>An example calo-image dataset used in the <a href="https://arxiv.org/abs/2205.09876">paper</a>: standard model.</p> <p>Each shard is named `calo-image_sm_${DATASET}_${INDEX}.tar.gz`. Each contains one data file. DATASET is in {train,test,private_test} to label the three independent splits for training, validation, and testing respectively. INDEX labels separate batches which should be trivially combined.</p> <p>Each data file is in <a href="https://www.h5py.org/">h5</a> format. Its data are under the key &quot;entries&quot; as an array with shape (n, 32, 32) corresponding to (event_index, eta, phi) for the n calorimeter images in an unrolled eta--phi surface.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Standardized Dataset of the Ecosystem's Water Use Efficiency, Gross Primary Productivity and the Evapotranspiration Deficit Index for 1982–2017 over the Middle East

<p>This data aimed to investigate the spatial-temporal variability of&nbsp;Standardized Actual Evapotranspiration (sAET), Gross Primary Productivity (sGPP) and Water Use&nbsp;Efficiency&nbsp;(WUE) anomalies series,&nbsp;and the Standardized Evapotranspiration Deficit Index (SEDI). The Middle East (ME),&nbsp;was selected as a case study to monitoring &nbsp;drought events as one of the major natural disasters for the ecosystem. To this end, the yearly gross primary production of GLASS, GIMMS, &nbsp;FloxCom, and VPM datasets for the study area spanning 1982&ndash;2017 was used to develop&nbsp;the sGPPR data. On the other hand, the Global Land Evaporation Amsterdam Model (GLEAM-version (v3.3a)), which estimated the several components of terrestrial evaporation (annual actual and potential evaporation (AET, PET)) was used for the same period this aimed to detect the variability of the SEDI.<br> This version of the yearly GLASS-sGPPR dataset (1982&ndash;2017) is available for the ME at 0.05&deg; spatial resolution, as the original data of &nbsp;the GPP-GLASS products, While, sGPPR dataset of GIMMS, &nbsp;FloxCom, and VPM are also at annual temporal resolution, and at 0.5 degree spatial resolution spanning 1982&ndash;2016 for GIMMS, &nbsp;FloxCom, and 2000-2016 for VPM (Excel wrokbook .xlsx). The SEDI data are also available at 0.25 degree spatial resolution for 1980&ndash;2018 ( Raster files (TIFF)). For more details about Standardization of the GPP and evapotranspiration deficit &nbsp;data see: <strong>Alsafadi, K., Al-Ansari, N., Mokhtar, A., Mohammed, S., Elbeltagi, A., Sammen, S. S., &amp; Bi, S. (2021). An evapotranspiration deficit-based drought index to detect variability of terrestrial carbon productivity in the Middle East. <em>Environmental Research Letters</em>.&nbsp;<a href="http://dx.doi.org/10.1088/1748-9326/ac4765">10.1088/1748-9326/ac4765</a></strong></p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Data from: Global Spore Sampling Project: A global, standardized dataset of airborne fungal DNA

<p><span>Novel methods for sampling and characterizing biodiversity hold great promise for re-evaluating patterns of life across the planet. The sampling of airborne spores with a cyclone sampler, and the sequencing of their DNA, have been suggested as an efficient and well-calibrated tool for surveying fungal diversity across various environments. Here we present data originating from the Global Spore Sampling Project, comprising 2,768 samples collected during two years at 47 outdoor locations across the world. Each sample represents fungal DNA extracted from 24 m<sup>3</sup> of air. We applied a conservative bioinformatics pipeline that filtered out sequences that did not show strong evidence of representing a fungal species. The pipeline yielded 27,954 species-level operational taxonomic units (OTUs). Each OTU is accompanied by a probabilistic taxonomic classification, validated through comparison with expert evaluations. To examine the potential of the data for ecological analyses, we partitioned the variation in species distributions into spatial and seasonal components, showing a strong effect of the annual mean temperature on community composition.</span></p> <p><span>The database is organized in five datasets in a csv format (columns separated by commas): (1) metadata providing the location, date, and time for each sample, along with sequencing depth and other essential information (metadata.csv); (2) species-level OTU tables per sample describing the number of sequences assigned to each species (otu.table.csv 3); (3) taxonomic classification of each species-level OTU (taxonomy.csv); (4) closest matching sequences and their taxonomy for ASVs in putatively fungal pseudophyla, which are included in (2) and (3) (fungi_pseudophyla.csv); and (5) closest matching sequences and their taxonomy for ASVs in putatively non-fungal pseudophyla, which are not included in the other datasets (nonfungi_pseudophyla.csv). The first four datasets can be linked to each other using the unique sample codes and the unique identifiers for species-level OTUs. </span><span>The three first datafiles are also provided in allData.RData which can be read into R as load("allData.RData").</span></p>

opencc-by-4.0May 2024View details →
zenodo40/100

JAZZVAR: A Dataset of Variations found within Solo Piano Performances of Jazz Standards for Music Overpainting

<p>Release of the MIDI data pairs that constitute the JAZZVAR dataset. See below for the abstract of the publication.</p> <p>The data is also available transposed to C/Am and subsequently, to all keys, with accompanying metadata.</p> <p>Abstract:</p> <p>Jazz pianists often uniquely interpret jazz standards. Passages from these interpretations can be viewed as sections of variation. We manually extracted such variations from solo jazz piano performances. The JAZZVAR dataset is a collection of 502 pairs of Variation and Original MIDI segments. Each Variation in the dataset is accompanied by a corresponding Original segment containing the melody and chords from the original jazz standard. Our approach differs from many existing jazz datasets in the music information retrieval (MIR) community, which often focus on improvisation sections within jazz performances. In this paper, we outline the curation process for obtaining and sorting the repertoire, the pipeline for creating the Original and Variation pairs, and our analysis of the dataset. We also introduce a new generative music task, Music Overpainting, and present a baseline Transformer model trained on the JAZZVAR dataset for this task. Other potential applications of our dataset include expressive performance analysis and performer identification.</p>

opencc-by-nc-sa-2.0May 2024View details →
zenodo40/100

Dataset: Standard BioTools Inc. (LAB) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Datasets on approved and ongoing standards for Cloud, Edge and IoT computing in the continuum and analysis and assessment of relevance

<p>Two datasets:</p> <ol> <li><span>the database of standards relevant to the Cloud-Edge-IoT continuum and to the ACES-EDGE Research and Innovation Action funded under the grant agreement No. 101093126 call HORIZON-CL4-2022-DATA-01-02.</span></li> <li><span>the Excel workbook with the analysis of the assessment of the standards in relation to the needs of the ACES-EDGE implementation.</span></li> </ol> <p><span>Both datasets will be used for a more deep assessment of the standardisation requirements of the ongoing technolgical developments.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Dataset for "Atmospheric CFC-11 and CCl4: a Free Calibration Standard for PTR-MS"

<p>Dataset for the publication: Not&oslash; and Holzinger (2024), &ldquo;Atmospheric CFC-11 and CCl4: a Free Calibration Standard for PTR-MS&rdquo;,&nbsp;<br><a title="Atmospheric CFC-11 and CCl4: a Free Calibration Standard for PTR-MS" href="https://doi.org/10.1016/j.ijms.2024.117311">https://doi.org/10.1016/j.ijms.2024.117311</a></p> <p>The data consists of raw data files of measurements and the processing code to calculate pseudo reaction rate constants of CFC-11 and CCl4 with H3O+.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Dataset to: Realistic accelerated stress tests for PEM fuel cells: Test procedure development based on standardized automotive driving cycles

<p>This is the dataset to the published article "Realistic accelerated stress tests for PEM fuel cells: Test procedure development based on standardized automotive driving cycles" (DOI: 10.1016/j.ijhydene.2023.08.292) in which the degradation of two commercial PEM fuel cell stacks was analyzed.&nbsp;<strong>Please cite this publication if you use the dataset in a publication as follows</strong>:</p> <p>P. Thiele, Y. Yang, S. Dirkes, M. Wick, S. Pischinger, Realistic accelerated stress tests for PEM fuel cells: Test procedure development based on standardized automotive driving cycles, Int. J. Hydrogen Energy 52 (Part D) (2024) 1065&ndash;1080, https://doi.org/10.1016/j.ijhydene.2023.08.292.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

A Compiled Archaeobiological Dataset for Central Asia's Chalcolithic through Bronze Age: macrobotanical and zooarchaeological data transcribed, standardized, and summarized from original publications

<p>This dataset contains archaeobotnaical and zooarchaeological data that have been compiled, transcribed, standardized, and summarized from original published data sources. Original data publications are given herein as a List of References (Microsoft Word file). These publications appeared between 1960-2022, presenting data in various formats, in various languages, and in scientific works that included journals, books, and conference proceedings - .</p> <p>Data have been compiled and are given in two Microsoft Excel files, one corresponding to archaeobotanical data and one to zooarchaeological data. On the first tab (worksheet) of each of these files, original publication sources are given in a summarized reference (Author, Year, Table/Figure Number) that corresponds to the full bibliographic reference in the accompanying List of References file. This first tab (worksheet) also summarizes additional information on the archaeological context of each dataset, collection and analysis methods (when reported), and the availabilty of other relevant and/or corresponding datasets. The remaining tabs (worksheets) in each file, organized alphabetically by author last name, tabularize the original data in a standardized format; these data are compiled and transcribed as necessary from the various formatting of original data sources, though they keep the original reported species names, table ordering, and numerical data. In the cases where totals were obviously erroneous or superceded by later or additional analyses, transcription notes have been offered directly in the worksheet.</p> <p>A fair number of these data sources are now out of print, have no digital distribution, and are otherwise difficult to access physically and/or linguistically. Accordingly, the sole aim in compiling these data here is to facilitate their widespread availablity, proper citation, and increased use within the international community of archaeological scientists working in Central Asia in the present day. Other scholars are encouraged to utilize these compiled datasets for foundational regional data and extended analyses, and to add to and expand these datasets going forward.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

D-PLACE dataset derived from Murdock and White 1969 'Standard Cross-Cultural Sample'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Murdock GP &amp; White DR. 1969. Standard Cross-Cultural Sample. Ethnology. 9:329–369.</p> </blockquote>

opencc-by-nc-4.0Nov 2023View details →
zenodo40/100

Dataset related to article "On the extension of the use of a standard operating procedure for nicotine, glycerol and propylene glycol analysis in e-liquids using mass spectrometry"

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo40/100

Dataset for the paper "Yield Performance of Standard Multicrystalline, Monocrystalline, and Cast-Mono Modules in Outdoor Conditions"

<p>Dataset for the paper "Yield Performance of Standard Multicrystalline, Monocrystalline, and Cast-Mono Modules in Outdoor Conditions", published at Energies.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

LAGOS-AND: A Large Gold Standard Dataset for MAG/OpenAlex Author Name Disambiguation

<p>We present a large gold standard dataset for author name disambiguation (AND) research (LAGOS-AND), which contains two sub-datasets, LAGOS-AND-BLOCK and LAGOS-AND-PAIRWISE. The datasets were automatically built by using the two authoritative sources, ORCID and DOI, based on the ORCID open database and an open literature database (MAG or OpenAlex).</p> <p>The currently available versions of the LAGOS-AND datasets are:</p> <ol> <li><strong>Version 1.0 (MAG+ORCID)</strong>: This is the initial version of the LAGOS-AND dataset, the evaluation results and quality control measures of this version dataset can be found in this paper <a href="https://arxiv.org/abs/2104.01821">https://arxiv.org/abs/2104.01821</a>.</li> <li><strong>Version 2.0 (OpenAlex+ORCID)</strong>: This version builds on OpenAlex instead of MAG because MAG was discontinued on 31 December 2021, and OpenAlex not only positions itself as a drop-in replacement for MAG but also keeps evolving by aggregating academic resources from other repositories. In addition, the pairwise-based sub-dataset (LAGOS-AND-PAIRWISE v2.0) improves the accuracy of labeled authorship (class label) as compared to LAGOS-AND-PAIRWISE v1.0.</li> </ol> <p>Note that there are other versions of the dataset, which we call &quot;pre-release&quot;. We recommend users to use the normal version of the dataset. These pre-release versions were originally intended to be released as the normal versions. However, during the preparation of the research paper, we found an issue with the dataset and the reviewers also made some reasonable requests for the dataset. This led us to update the dataset. Unfortunately, the Zenodo platform does not allow updates to the same version of dataset, so we had to create new versions. The created pre-release datasets are as follows:</p> <ol> <li><strong>Version 2.0-alpha</strong>: For few samples in LAGOS-AND-PAIRWISE, the class labels are incorrect. We improve the accuracy of the class label in the normal Version 2.0 dataset by using a better random sampling approach.</li> <li><strong>Version 1.0-beta</strong>: We created this version because a sub-dataset of this version LAGOS-AND-PAIRWISE contains only ~500K author pairs, while it should contain ~1M author pairs, as described in our paper. We fixed the problem in the Version 1.0 dataset.</li> <li><strong>Version 1.0-alpha</strong>: The earliest dataset uploaded to Zenodo, corresponding to the original dataset before addressing the reviewers&#39; comments and suggestions. In contrast, the dataset in Version 1.0 is the dataset after the reviewers&#39; comments and suggestions have been addressed.</li> </ol>

opencc-by-4.0Feb 2021View details →
zenodo40/100

Dataset for the publication "The Use of Voltage Transformers for the Measurement of Power System Subharmonics in Compliance With International Standards"

<p>This is dataset for paper published:</p> <p>G. Crotti, G. D&rsquo;Avanzo, P. S. Letizia and M. Luiso, &quot;The Use of Voltage Transformers for the Measurement of Power System Subharmonics in Compliance With International Standards,&quot; in&nbsp;<em>IEEE Transactions on Instrumentation and Measurement</em>, vol. 71, pp. 1-12, 2022, Art no. 9005912, doi: 10.1109/TIM.2022.3204318.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Datasets for "Relic gravitational waves from the chiral plasma instability in the standard cosmological model"

<pre>This directory contains an index.html file with links to the run directories and idl plotting routines with secondary data for the other figures for the paper &quot;Relic Gravitational Waves from the Chiral Plasma Instability in the Standard Cosmological Model&quot;. If anything turns out to be incomplete, please email brandenb@nordita.org.</pre>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Silvicultural and economic dataset for even-aged and coppice-with-standards managements in France.

<p>Data set from&nbsp;DEFIFORBOIS project &quot;Development and sustainability of the forest-wood sector in the Centre region&quot;, PSDR 4 project Centre-Val de Loire Region, France. This study has been carried out also with financial support from the French National Research Agency (ANR) in the frame of the Investments for the future Programme, within the Cluster of Excellence COTE (ANR-10-LABX-45) through the Project LUCAS.</p>

opencc-by-4.0Sep 2020View details →
zenodo36/100

STAVER: A Standardized Benchmark Dataset-Based Algorithm for Effective Variation Reduction in Large-Scale DIA-MS Data

<p>This project focuses on developing and applying STAVER, an innovative DIA algorithm designed to eliminate non-biological noise and variability from the large-scale DIA-MS study dataset analyses. STAVER is a flexible framework that utilizes prior knowledge regarding peptide separation coordinates (RT) and fragment ion intensities from the standard benchmark datasets, which effectively mitigates non-biological noise potential during library searches, enhancing spectrum identifications and protein quantification accuracy. Furthermore, the robustness and broad applicability of STAVER were validated in multiple large-scale DIA datasets from different platforms and laboratories, demonstrating significantly improved precision and reproducibility of protein quantification. It facilitates the comparative and integrative analysis of DIA datasets across different platforms and laboratories, enhancing the consistency and reliability of findings in clinical research. The project aims to promote the adoption of hybrid library search and improve the sensitivity and quality of DIA proteomics data through the open-source STAVER software package.</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Biaxial seismic response of base-column connections in sub-standard steel buildings: dataset

<p><span>A complete dataset on the response of (semi-rigid and partial strength) column base plate connections in a substandard steel frame tested experimentally, are provided. Free-vibration, cyclic and pseudo-dynamic tests were carried out </span><span>at the Structures Laboratory (STRULAB) of the University of Patras in the framework of </span><span>H2020 EU-funded "</span><span>Engineering Research Infrastructures for European Synergies (ERIES)" project.</span></p> <p><span>&nbsp;</span></p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Dataset: Improving the prediction of fertilizer phosphorus availability to plants with simple, but non-standardized extraction techniques

<p>Dataset: Improving the prediction of fertilizer phosphorus availability to plants with simple, but non-standardized extraction techniques</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record