Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,481 results for “data set”

Learn how ShareScore rates datasets ↗
zenodo36/100

Data Set for Predicting the Performance of ATL Model Transformations

<p>Model transformation languages are special-purpose languages, which are designed to define transformations as comfortably as possible, i.e., often in a declarative way. With the increasing use of transformations in various domains, the complexity and size of input models are also increasing. However, developers often lack suitable models for performance testing. We have therefore conducted experiments in which we predict the performance of model transformations based on characteristics of input models using machine learning approaches. This dataset contains our raw and processed input data, the scripts necessary to repeat our experiments, and the results we obtained.</p> <p>Our input data consists of the time measurements for six different transformations defined in the Atlas Transformation Language (ATL), as well as the collected characteristics of the real-world input models that were transformed. We provide the script that implements our experiments. We predict the execution time of ATL transformations using the machine learning approaches linear regression, random forests and support vector regression using a radial basis function kernel. We also investigate different sets of characteristics of input models as input for the machine learning approaches. These are described in detail in the provided documentation.pdf. The results of the experiments are provided as raw data in individual cvs files. Additionally, we calculated the mean absolute percentage error in % and the 95th percentile of the absolute percentage error in % for each experiment and provide these results. Furthermore, we provide our Eclipse plugin, which collects the characteristics for a set of given models, the Java projects used to measure the execution time of the transformations, and other supporting scripts, e.g. for the analysis of the results.</p> <p>A short introduction with a quick start guide can be found in README.md and a detailed documentation in documentaion.pdf.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Raw data and stimuli for assessing the perceived reverberation in different rooms for a set of musical instrument sounds

<p>This set of data and sound stimuli was used in the study by Osses, McLachlan, and Kohlrausch (2020) to assess the perceived reverberation --including measurements and simulations-- for different instrument sounds in eight different rooms. The following are the directories that are provided:</p> <ul> <li><strong>00-Experiment-WAE_GM_201712</strong>: Web Audio Evaluation tool (WAE) used to run the listening experiment with 24 participants. Follow the instructions in README.txt to get the experiment running.</li> <li><strong>01-Stimuli</strong>: Sound stimuli as exactly used during the listening experiments.</li> <li><strong>02-Raw-data</strong> and <strong>03-Results-summary</strong>: Outputs from WAE for each of the participants. The raw data contained in these XML files were extracted and stored in &#39;03-Results-summary&#39;</li> <li><strong>04-Stimuli-9s-for-simulations</strong>: Same sounds as in &#39;01-Stimuli&#39; but truncated to have a duration of 9 s. These sounds were used as input to an implementation (Osses et al. 2017, 2020) of the model by van Dorp et al. (2013).</li> </ul> <p>The paper figures can be reproduced in two MATLAB toolboxes: fastACI (script: publ_osses2020a_JASA_EL_figs.m) and AMT (script: exp_osses2020.m, availability as of 2023).</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

WorldCereal open global harmonized reference data repository (CC-BY-NC licensed data sets)

<p>Within the <strong>ESA funded</strong> WorldCereal project we have built an open harmonized reference data repository at global extent&nbsp;for model training or product validation&nbsp;in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then&nbsp;harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes&nbsp;(LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication&nbsp;includes those harmonized&nbsp;data sets of which the original data set was&nbsp;published under the CC-BY-NC license or a license similar to CC-BY-NC. See document &quot;_In-situ-data-World-Cereal - license - CC-BY-NC.pdf&quot; for an overview of the original data sets. Currently this publication only includes a few small data sets for Tanzania originating from a disease monitoring program of the International Maize and Wheat Improvement Center (CIMMYT). CIMMYT made more data available for&nbsp;countries like Kenya, Ethiopia, Rwanda and, Malawi. However due project contraints these data sets were not yet harmonized.</p>

opencc-by-nc-4.0Dec 2022View details →
zenodo36/100

Data set for: Quantifying internal stress and demagnetization effects for natural multidomain magnetite and magnetite-ilmenite intergrowths

<p>This data set contains the raw measurement data used for the manuscript &quot;<em><strong>Quantifying internal stress and demagnetization effects for natural multidomain magnetite and magnetite-ilmenite intergrowths</strong></em>&quot;, submitted to the <em>Journal of geophysical research - Solid Earth</em>.&nbsp;</p> <p>The data set includes:&nbsp;</p> <ul> <li>MPMS_HighTemperature_SRW</li> <li>VSM_HighTemperature_SRW</li> <li>VSM_FORC</li> <li>VSM_HighTemperature_TC</li> <li>VSM_Orientation</li> <li>VSM_Preisach</li> </ul> <p>All VSM data files are in the format: MicroMag 2900/3900 Data Files (Series 0016.002).&nbsp;Both the VSM data files and the MPMS data files provide specimen information. Each file contains information on the instrument, settings, measurement, script, and the measured data.&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

data set regarding to article Comparison of diagnostic value of SARC-F and its four modified versions in Polish community-dwelling older adults

<p>This data set corresponds with the article titled Comparison of diagnostic value of SARC-F and its four modified versions in Polish community-dwelling older adults</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Efficacy of low temperature plasma in eliminating fungi and aflatoxins - data set

<p>Set of data generated during the experimental runs</p>

opencc-by-4.0Feb 2023View details →
dryad36/100

Consilience across multiple, independent genomic data sets reveals species in a complex with limited phenotypic variation

<p>Species delimitation in the genomic era has focused predominantly on the application of multiple analytical methodologies to a single massive parallel sequencing (MPS) data set, rather than leveraging the unique but complementary insights provided by different classes of MPS data. In this study we demonstrate how the use of two independent MPS data sets, a sequence capture data set and a single nucleotide polymorphism (SNP) data set generated via genotyping-by-sequencing, enables the resolution of species in three complexes belonging to the grass genus <em>Ehrharta, </em>whose strong population structure and subtle morphological variation limit the effectiveness of traditional species delimitation approaches. Sequence capture data are used to construct a comprehensive phylogenetic tree of <em>Ehrharta </em>and to resolve population relationships within the focal clades, while SNP data are used to detect patterns of gene pool sharing across populations, using a novel approach that visualises multiple values of K. Given that the two genomic data sets are fully independent, the strong congruence in the clusters they resolve provides powerful ratification of species boundaries in all three complexes studied. Our approach is also able to resolve a number of single-population species and a probable hybrid species, both which would be difficult to detect and characterize using a single MPS data set. Overall, the data reveal the existence of 11 and five species in the <em>E. setacea</em> and <em>E. rehmannii </em>complexes, with the <em>E. ramosa</em> complex requiring further sampling before species limits are finalized. Despite phenotypic differentiation being generally subtle, true crypsis is limited to just a few species pairs and triplets. We conclude that, in the absence of strong morphological differentiation, the use of multiple, independent genomic data sets is necessary in order to provide the cross-data set corroboration that is foundational to an integrative taxonomic approach.</p>

opencc-zeroFeb 2023View details →
zenodo36/100

data set for Design of architectural environment integration of cyber-physical systems based on image segmentation and retrieval technology

<p>This data set is used to implement the project&nbsp;&nbsp;Design of architectural environment integration of cyber-physical systems based on image segmentation and retrieval technology</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Training, Validation and Test Sets for paper 'A Little Data goes a Long Way: Automating Seismic Phase Arrival Picking at Nabro Volcano with Transfer Learning'

<p>Training, Validation and Test Data for model presented in&nbsp;paper &#39;A Little Data Goes A Long Way: Automating Seismic Phase Arrival Picking at Nabro Volcano with Transfer Learning&#39;, submitted to Journal of Geophysical Research: Solid Earth.</p> <p>Files:</p> <p>- train_events_2498.h5 = training set of seismic waveforms (events with P-/S-wave labelled arrivals only, i.e., no noise waveforms)</p> <p>- train_events_2498.pkl = event training set metadata (UTC P-/S-wave phase arrival times)</p> <p>- train_noise_2498.h5 = training set of seismic waveforms (noise sections only, i.e., no event waveforms)</p> <p>- train_noise_2498.pkl = noise training set metadata (UTC time&nbsp;for training noise waveforms)</p> <p>- val_events.h5 = validation set of seismic waveforms (events with P-/S-wave labelled arrivals only, i.e., no noise waveforms)</p> <p>- val_events.pkl = event validation set metadata (UTC P-/S-wave phase arrival times)</p> <p>- val_noise.h5 = validation&nbsp;set of seismic waveforms (noise sections only, i.e., no event waveforms)</p> <p>- val_noise.pkl = noise validation set metadata (UTC time&nbsp;for validation noise waveforms)</p> <p>- test.h5 = test&nbsp;set of seismic waveforms (events and noise)</p> <p>- test_events.pkl = event test set metadata (UTC P-/S-wave phase arrival times for test event waveforms)</p> <p>- test_noise.pkl = noise test set metadata (UTC time for test noise waveforms)</p> <p>- nabro_2011-247.mseed = 24 hours seismic data from Nabro Urgency Array (2011-09-04), saved in mseed format (e.g., can be read with obspy)</p> <p>- nabro_2011-269.mseed = 24 hours seismic data from Nabro Urgency Array (2011-09-26), saved in mseed format (e.g., can be read with obspy)</p> <p>&nbsp;</p> <p>Further details and code for reading and using&nbsp;these files can be found at the GitHub repo for this paper:&nbsp;<a href="https://github.com/sachalapins/U-GPD">https://github.com/sachalapins/U-GPD</a></p> <p>&nbsp;</p>

opencc-by-4.0Feb 2021View details →
zenodo36/100

Gephyromantis (subgenus Gephyromantis) revision data sets

<p>Original sound recordings used for bioacoustic analysis as well as DNA sequence alignments for a taxonomic revision of the subgenus Gephyromantis (genus Gephyromantis): G. boulengeri and G. blanci species complexes.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Data Set to doi.org/10.3390/app13042747 "Robot Assisted THz Imaging with a Time Domain Spectrometer"

<p>Data set to article &quot;Robot Assisted THz Imaging with a Time Domain Spectrometer&quot; in Applied Sciences 2023, 13, 2747.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Data set for the article "One-time freeze-thawing or carbon input events have long-term legacies in soil microbial communities"

<p>The following are data and code used for statistical analysis and figure plotting in the manuscript</p> <p>Gorka et al. (2023) &quot;One-time freeze-thawing or carbon input events have long-term legacies in soil microbial communities&quot;, Geoderma, <a href="https://doi.org/10.1016/j.geoderma.2023.116399">https://doi.org/10.1016/j.geoderma.2023.116399</a></p> <p>It contains the following files:</p> <ol> <li>Microbial PLFA and NLFA analysis files (contained in <em>Fatty_acid_analysis.zip</em>) <ul> <li>PLFA and NLFA abundance data, in nmol C g<sup>-1</sup> dw (in <em>fatty_acid_data.csv</em>)</li> <li>Taxonomic specificities of fatty acids, needed for R-script to run (in <em>phylum.csv</em>)</li> <li>An R Script reproducing the statistical analysis, figure plotting and output for tables, as used in the manuscript (<em>Fatty_acid_analysis.R</em>)</li> </ul> </li> <li>Soil C and N stoichiometry analysis files (contained in <em>TOC_TN_analysis.zip</em>) <ul> <li>Dissolved organic C (DOC), total dissolved N (TN), and C and N in microbial biomass (Cmic, Nmic) abundance data, in mg g<sup>-1</sup> dw (in <em>toc_tn_data.csv</em>)</li> <li>An R Script reproducing the statistical analysis, figure plotting and output for tables, as used in the manuscript (<em>TOC_TN_analysis.R</em>)</li> </ul> </li> </ol>

opencc-by-4.0Feb 2023View details →
zenodo36/100

NEXAFS images and raw data set used for a paper by Gonçalves Jr. et al.

<p>NEXAFS images (regions&nbsp;maps) and data set of aerosol particle compositions measured using synchrotron-based multi-element microscopic speciation of individual microparticles (STMX/NEXAFS - Scanning Transmission X-ray Microscopy with Near- Edge X-ray Absorption Fine Structure Spectroscopy combined)&nbsp;used in results at a paper by Gon&ccedil;alves Jr.&nbsp;et al. The samples were collected during the 2014 Summer Criosfera-1&nbsp;(West Antarctica) campaign.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Radiosonde and sun photometer data set

<p>Radiosonde and sun photometer data set used in the paper &quot;Precipitable water vapor retrievals using a ground-based infrared sky camera in subtropical South America&quot; - Atmospheric Measurement Techniques, by Hack et al. 2023.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Molecular and 13C isotopic composition of diacids and related compounds in wintertime PM2.5 at three sites over Northeast Asia – Data set

<p>To investigate the spatial changes in atmospheric organic aerosols (OA) loading and composition over the northeast Asian region, fine aerosols (PM<sub>2.5</sub>) were collected at three sites: Tianjin (TJ), North China and in Padori (PD) and Daejeon (DJ), South Korea, during winter 2019. We studied the molecular distributions and compound-specific carbon isotopic composition (&delta;<sup>13</sup>C) of dicarboxylic acids (diacids), oxocarboxylic acids (oxoacids) and &alpha;-dicarbonyls as well as the carbonaceous and inorganic ionic components in two sets of selected samples, representing a clean and a polluted period, during the campaign.&nbsp;Based on the molecular distributions and&nbsp;&delta;<sup>13</sup>C&nbsp;of diacids and related compounds and mass ratios and linear relations of selected species, we discuss the origins of OA and the influence of long-range transported air masses on their loading and composition over Northeast Asia.</p>

opencc-by-4.0Mar 2023View details →
dryad36/100

The HAInich: A multidisciplinary vision data-set for a better understanding of the forest ecosystem

<p>We present a multidisciplinary forest ecosystem 3D perception dataset. The dataset was collected in the Hainich-Dün region in central Germany, which includes two dedicated areas, which are part of the Biodiversity Exploratories – a long-term research platform for comparative and experimental biodiversity and ecosystem research. The dataset combines several disciplines, including computer science and robotics, biology, bio-geochemistry, and forestry science. We present results for common 3D perception tasks, including classification, depth estimation, localization, and path planning. We combine the full suite of modern perception sensors, including high-resolution fisheye cameras, 3D dense LiDAR, differential GPS, and an inertial measurement unit, with ecological metadata of the area, including stand age, diameter, exact 3D position, and species. The dataset consists of three hand-held measurement series taken from sensors mounted on a UAV during each of three seasons: winter, spring, and early summer. This enables new research opportunities for the forest ecosystem and paves the way for testing forest environment 3D perception tasks and mission set automation.</p>

opencc-zeroSep 2022View details →
zenodo36/100

BE-KONFORM data set (Group A) (Bedarfsermittlung im Rahmen der Erstellung des Konzepts für das Forschungsdatenmanagement an der Medizinischen Fakultät der Universität Freiburg

<p>This dataset was collected within the BE-KONFORM study investigating employee&#39;s needs regarding research data management at the Medical Faculty of the University of Freiburg. The full dataset captures 236 complete cases. The study included a randomized module allocating subjects to one of two groups. Group A received the information that data will be published (n=113) and group B did not. Due to data protection law, only data of group A, where written informed consent was given (n=112) could be published here. This dataset had to be prepared for publication in order to avoid de-anonymisation of subjects. Therefore, variable [anzahl_m] was recoded into a categorical variable. Open text answers, where combination of variables might lead to an identification of the subjects were changed, where necessary. Changes are marked with brackets: [changed text]. Information on survey mode and sampling is provided in the data note.</p>

openDec 2022View details →
zenodo36/100

Supporting Data Set for Paper "Using GUI Test Videos to Obtain Stakeholders' Feedback"

<p>This data set is a supporting material for an accepted paper &quot;Using GUI Test Videos to Obtain Stakeholders&rsquo; Feedback&quot; on <a href="https://conf.researchr.org/track/icssp-2023/">ICSSP 2023</a>.</p> <p>This dataset consists of</p> <ul> <li>Questionnaires of control and experimental groups <ul> <li> <p>Questionnaire-Control Group (German).pdf -&gt; Original quetionnaire for control group (in German)</p> </li> <li> <p>Questionnaire-Control Group (translated)-v03.pdf -&gt; Translated quetionnaire for control group (in English)</p> </li> <li> <p>Questionnaire-Experimental Group (German).pdf -&gt; Original quetionnaire for experimental group (in German)</p> </li> <li> <p>Questionnaire-Experimental Group (translated)-v03.pdf -&gt; Translated quetionnaire for experimental group (in English)</p> </li> </ul> </li> <li>Survey-Data-v25.ods -&gt; Collected data through questionnaires</li> <li>Calculate-Mann-Whitney-U-Test-v07.ods -&gt; Detailed calculation of Mann Whitney U Test</li> <li>The videos of the ten scenarios in this study are also available on OneDrive <a href="https://1drv.ms/f/s!AtqkJ5cB802BoABI7w_Bft9H8Psi?e=upToep">https://1drv.ms/f/s!AtqkJ5cB802BoABI7w_Bft9H8Psi?e=upToep</a> , where you can play them directly in browsers.</li> </ul>

opencc-by-4.0Mar 2023View details →
dryad36/100

Data for: Capturing complex interactions in disease ecology with simplicial sets

<p>Here we provide archived code for: Code for Capturing Complex Interactions in Disease Ecology with Simplicial Sets. Ecology Letters. The code provided can be used to generate the figures used in the manuscript as well as to generate appendix 3 in the supplementary materials. In the paper, we describe how higher-order network approaches can be applied in disease ecology research. We explain <em>what</em> simplicial sets are; <em>why</em> their use would be beneficial in different subject areas; <em>where</em> these areas are: social, transmission, movement/spatial and ecological networks; and <em>when</em> using them would help most in each context. To demonstrate their application, we develop a novel approach to identify how pathogens persist within a host population (see code for Appendix 3 in this repository). We also provide an overview of <em>how</em> to use simplicial sets, highlighting specific metrics, generative models and software. Finally, we <em>synthesize</em> key research questions simplicial sets will help us answer and highlight the methodological developments required.</p>

opencc-zeroMar 2023View details →
zenodo36/100

Figures and Data Set of "The Effelsberg survey of FU Orionis and EX Lupi objects I. - Host environments of FUors/EXors traced by NH3"

<p>This is the full version of Figure 1 and the fits files for the NH<sub>3</sub> detections for the paper&nbsp;"The Effelsberg survey of FU Orionis and EX Lupi objects I. - Host environments of FUors/EXors traced by NH<sub>3</sub>".</p> <p>The paper is published in Astronomy &amp; Astrophysics, and can be found at: <a href="https://doi.org/10.1051/0004-6361/202244911">https://doi.org/10.1051/0004-6361/202244911</a></p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record