Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
41,236
datasets available to search
ShareScore release 0.9.0
Dataset results
41,236 results for “Review of reviews”
Supporting data for review article: The Global Distribution, Formation, and Fate of Mineral-Associated Soil Organic Matter Under a Changing Climate – A Trait-Based Perspective
<p>Supporting data and code for review article: Sokol N.W., Whalen E.D., Kallenbach C., Pett-Ridge J., Georgiou K. The Global Distribution, Formation, and Fate of Mineral-Associated Soil Organic Matter Under a Changing Climate – A Trait-Based Perspective. <em>Functional Ecology, </em>2022.</p> <p>We leveraged data from a global synthesis of soil fractionation measurements (DOI: 10.5281/zenodo.5987415). For this review article, we specifically focused on measurements of bulk and mineral-associated soil organic carbon concentrations (reported in units of gC/kg soil) and the proportion of bulk soil organic carbon that is mineral-associated (reported as a %). This subset also includes auxiliary data regarding climate and biome characteristics extracted from the synthesized papers; for more variables, see the original full dataset. Köppen-Geiger climate zones were extracted from a georeferenced global database (using R package 'kgc' v1.0.0.2) with site coordinates, where available. Three files are provided in this repository: (1) data file, (2) metadata file, and (3) code for manuscript figures and summary statistics.</p>
Data used to create figures and tables in the ACP manuscript "Two-way coupled meteorology and air quality models in Asia: a systematic review and meta-analysis of impacts of aerosol feedbacks on meteorology and air quality" by Gao et al. (2022)
<p>This dataset contains the original data that extracted from all collected papers refering applications of two-way coupled models in Asia. It is supplied to the review paper, which titled as "Review on two-way coupled meteorology and air quality models in Asia: impacts of aerosol feedbacks on meteorology and air quality". The dataset includes three excel files (in the format of xlsx) as follows:</p> <p>1. Basic information of literatures (Table S1.xlsx)</p> <p>2. Model performance metrics (Table S2.xlsx)</p> <p>3. Quantitative results of aerosol effects on meteorological and air quality variables (Table S3.xlsx)</p> <p>4. Basic information of model setup for two-way coupled model applications in Asia (Table S4.xlsx)</p> <p>5. Summary of aerosol-induced variations of simulated shortwave and longwave radiative forcing at the bottom and top of atmosphere and in the atmosphere in Asia (Table S5.xlsx)</p> <p>.</p>
Amazon product reviews (mock dataset)
<p>About</p> <p>This is a mock dataset with Amazon product reviews. Classes are structured: 6 "level 1" classes, 64 "level 2" classes, and 510 "level 3" classes.</p> <p><br> 3 files are shared:</p> <ul> <li>train_40k.csv - training 40k Amazon product reviews</li> <li>valid_10k.csv - 10k reviews left for validation</li> <li>unlabeled_150k.csv - raw 150k Amazon product reviews, these can be used for language model finetuning.</li> </ul> <p>Level 1 classes are: health personal care, toys games, beauty, pet supplies, baby products, and grocery gourmet food.</p> <p>Dataset originally from <a href="https://www.kaggle.com/datasets/kashnitsky/hierarchical-text-classification">https://www.kaggle.com/datasets/kashnitsky/hierarchical-text-classification</a></p>
Supplemental Material for a Systematic Literature Review on Benchmarks for Evaluating Debugging Approaches
<p>Bug benchmarks are used in development and evaluation of debugging approaches. Quantitative performance comparison of different debugging approaches is only possible when they have been evaluated on the same dataset or benchmark. However, benchmarks are often specialized towards usage for certain debugging approaches in their contained data, metrics, and artifacts. Such benchmarks can not be easily used on debugging approaches outside their scope as such approach may rely on specific data such as bug reports or code metrics not included in the dataset. Furthermore, benchmarks vary in their size w.r.t. the number of subject programs and the size of the individual subject programs. For these reasons, we have performed a systematic literature review where we have identified 73 benchmarks that can be used to evaluate debugging approaches.</p> <p>We compare the different benchmarks with respect to their size and the provided information such as bug reports, contained test cases, and other code metrics. Furthermore, we have investigated how well the benchmarks realize the <a href="https://www.go-fair.org/fair-principles/">FAIR guiding principles</a>. This comparison is intended to help researchers to quickly identify all suitable benchmarks for evaluating their specific debugging approaches. More information can be found in the publication:</p> <blockquote> <p>Thomas Hirsch and Birgit Hofer: "A Systematic Literature Review on Benchmarks for Evaluating Debugging Approaches", Journal of Systems and Software, in press, 2022.</p> </blockquote>
Raw data for the book chapter "Review of Haptic and Computerized (Simulation) Games on Climate Change"
<p>Raw data used for the book chapter "Gerber, A., Ulrich, M., Wäger, P. (2021). Review of Haptic and Computerized (Simulation) Games on Climate Change. In: Wardaszko, M., Meijer, S., Lukosch, H., Kanegae, H., Kriz, W.C., Grzybowska-Brzezińska, M. (eds) Simulation Gaming Through Times and Disciplines. ISAGA 2019. Lecture Notes in Computer Science(), vol 11988. Springer, Cham. <a href="https://doi.org/10.1007/978-3-030-72132-9_24">https://doi.org/10.1007/978-3-030-72132-9_24</a>"</p> <p>The documents include the raw data (both as .csv and .xlsx files with the same content), as well as the publication (.pdf file). The data collection process and the data itself are described in the publication. The data is published as "supplementary material" on the publisher's homepage.</p>
Social Innovations for Circularity in the Built Environment – a Scoping Review and Classification. Supplementary Data to the Bibliometric Review.
<p>To transition the built environment (BE) towards circularity, i.e., maximizing the time resources spend in the BE, thus minimizing negative environmental impacts of resource usage, social innovations (SI) – understood as new ways of doing, organizing, framing, and knowing – are just as important as technological advancements. This article provides an overview of the state of the knowledge regarding SI that contribute to circularity in the BE, and proposes a framework for coherently classifying such SI in terms of their main categories and effects.</p> <p>To identify and understand the current knowledge regarding social innovations (SI) that contribute to circularity in the built environment (BE), a bibliometric review of scientific literature is conducted. It shows that the term social innovation is not frequently used in this contexts although the buzzwords circularity and circular economy are themselves often framed as SI. To assess the characteristics and contribution of SI to circularity in the BE, a scoping review that includes grey literature into the context was conducted and a framework developed to classify predominant SI using concepts from transition studies as well as the systems thinking approach. The framework is designed to help assess the potential of SI and to identify research and/or action gaps. Key findings are (1) There is a broad diversity of SIs that contribute to circularity in the BE already in the focus of research, although they are not always identified as such; (2) Most SI focus on either the design or the demolition phase (i.e., market related phases), whereas user-centered SI are less frequently discussed; and (3) It is crucial to also consider potential sustainability goal conflicts in order to guide policies that address SI as a solution.</p> <p>These datasets are the basis to the bibliometric literature review.</p>
How Developers Review Tests in GitHub?
<p>This dataset contains 330 code reviews with 40 tips, 16 request categories, 8 response categories, and 13 Pull Request and comment features. This dataset's column titles are:</p> <p>Project Owner Project Repo Project URI Project Language Pull Request URI Pull Request ID Pull Request Author Pull Request Merge Commit Hash Review comment URI Review comment ID Review comment Text Solving Commit Hash Solving Commit URI Validation Request Category Tips Response Category Test Case Unspecified Test Method Fix SUT Optional Test Improve Test Refactor Test Test Class Fix Test Test Branch Test Statement Achieve Specific Coverage Goal ML Model Test Prevent Exploit Disagreement Revert Test Removal Unit Test Code Snippet End-to-End Test Edge Case Expected Exception Filepath Set Up Parametric Negative Test Test Double Type Support Based On External Resource Integration Test Positive Test Avoid Wrong API Usage Increase Testability Rename Test Dependency Reproduce Issue More Specific Test Event Test Assert Message Fix Based On Test Boundary Test Regression Test Async Thread Consistency Test Move Test Compilation Check Extract Member Cache Modifier Empty Test File Merge Test Readability Sleep Invalid Test Remove Reflection</p>
Variation in Detected Adverse Events using Trigger Tools: A Systematic Review and Meta-Analysis
<p>Raw data sets for the meta-analysis.</p> <p>Data collection file with all the information extracted from the included studies.</p> <p>QAT file with the information from the quality assessment tool (QAT) for all included studies.</p> <p>ReadMe with information on data sets and updates.</p> <p>Codebooks for both data sets.</p>
A Review Of Metaheuristics in Fuzzy Time Series Applied To Zero Inflated Datasets
<p>The manufacturing efficiency reflects directly on the use of natural resources and leads to a higher environmental impact than needed. Efficiency of an industry can be achieved in many ways, but it always starts with demand management. However some products have erratic and irregular demand patterns as the nature of the usage varies, and this often leads to zero inflated demand datasets, said datasets are difficult to forecast due to the nature of traditional models which usually use moving averages, state of the art machine learning models can achieve good results but use too much data for training. Under this background, this paper investigates the Fuzzy Time Series models and how it evolved from its inception to present time and how the usage of metaheuristics can help with forecasting demand on a small dataset with a high count of zeros, then applies the techniques to other zero inflated dataset to verify its generalization capabilities. Finally another model is applied as comparison.</p>
A Systematic Review on Techniques and Approaches to Estimate Mobile Software Energy Consumption (SUSCOM Dataset)
<p>Dataset and replication data for the systematic review entitled "A Systematic Review on Techniques and Approaches \\to Estimate Mobile Software Energy Consumption".</p>
GERDAT011 Literature search for publication - Geriatric assessment in the management of older patients with cancer – a systematic review (update).xlsx
<p>Search data belonging to the publication Geriatric assessment in the management of older patients with cancer – a systematic review (update)</p>
Philosphical papers reviewed and scored for their compliance with a vegan ethic 1975-2020
<p>This is the updated dataset to complement the submitted manuscript "<strong>Has the philosophical case for animal liberation been proved? A systematic and narrative review of the philosophical literature 1975-2020."</strong></p>
Privacy-by-Design Maturity Model: literature review, coding, model creation and evaluation
<p>Results from two multivocal literature reviews (MLRs) and subsequent coding, formulation of capabilities and dependencies, creation of maturity matrix and evaluation results. Used in the creation of a PbD domain model and extraction of core activities for PbD in Information Systems design. Part of the <a href="https://www.privacymaturity.org/" target="_blank" rel="noopener">Privacy-by-Design Maturity</a> research project by the <a href="https://www.uu.nl/en/research/ai-labs/ai-lab-for-the-public-services" target="_blank" rel="noopener">AI Lab for Public Services</a>.</p>
Classification and frequency of climate change drivers and responses of small-scale fishers found in literature review
<p>Climate change hazards were classified into resource availability and fishing operations or both following the framework proposed by Cheung et al (2012). Response units were firstly classified into overarching responses and then categorized as suggested by the adaptive-transformative framework of Barnes et al.<sup> </sup>(2020). Adaptation units that did not represent an active adaptation response were classified as remaining. More than one hazard could be attributed to each fishers' response.</p>
Quantification of ADHD Medication in Biological Fluids with Liquid Chromatography: A Comprehensive Review - Metadata
<p>This file is the metadata related to the publication "Quantification of ADHD Medication in Biological Fluids with Liquid Chromatography: A Comprehensive Review".</p>
Digital government HRM literature review
<p>Digital government HRM index</p> <p> </p> <p>Data collected and processed as part of the ODDEA (Overcoming Digital Divide Between Europe and Southeast Asia) EU research project (<em>Project ID: HORIZON MSCA-SE 101086381)</em></p>
2023 Utrecht University Open Access Monitor (peer reviewed journal articles)
<p>Results of the OA monitor of Utrecht University (UU) and University Medical Center Utrecht(UMCU) for the year 2023. It lists the open access availability of all peer reviewed journal articles registered in the CRIS (Pure) of Utrecht University and/or University Medical Center Utrecht. </p>
University dropout: A systematic review of the main determinant factors
<p><strong><span>Introduction:</span></strong><span> This research is a systematic review aimed at synthesizing scientific evidence on the causes of university dropout, focusing on the subcategories of vocational guidance, academic performance, socioeconomic status, and institutional aspects between 2020 and June 2024. <strong>Methods:</strong> Only articles addressing university dropout were considered, analyzing dimensions such as vocational guidance, academic performance, socioeconomic status, and institutional aspects. Articles published in indexed scientific journals with double-blind, double-blind peer, or open reviews between 2020 and June 2024 were included. The main databases used were Scopus, Web of Science, and Google Scholar. To assess the risk of bias in qualitative studies, the criteria from the article "Validity criteria for qualitative research: three epistemological strands for the same purpose" were used. For quantitative studies, the criteria from the article "Evaluating survey research in articles published in Library Science journals" were followed. For mixed-method studies, both sets of criteria were combined. <strong>Results:</strong> A total of 23 studies were included: 15 quantitative (65.22%), 3 qualitative (13.04%), and 5 mixed-method (21.74%). All studies (100%) addressed the subcategories of socioeconomic status and institutional aspects. Regarding the academic performance subcategory, 86% of the studies addressed it, while the vocational guidance subcategory was covered by 73.91% of the studies. <strong>Conclusions:</strong> Vocational guidance, academic performance, socioeconomic status, and institutional aspects are crucial for reducing university dropout. Providing adequate professional guidance, academic support, financial assistance, and strong institutional support is fundamental to improving student retention and academic success.</span></p>
Open Access levels of Dutch universities' output 2016-2017 (articles & reviews): green, gold, hybrid and bronze - May 2018
<p>Using Web of Science and Unpaywall data, we here provide an update of Open Access (OA) levels of Dutch universities, for 2016 and 2017.</p> <p>Our previous analysis (<a href="http://doi.org/10.5281/zenodo.1133759">10.5281/zenodo.1133759</a> and <a href="http://doi.org/10.7287/peerj.preprints.3520v1">10.7287/peerj.preprints.3520v1</a>) looked at OA classification as included in Web of Science (gold and green OA, based on Unpaywall data), and supplemented that with a breakdown of gold OA into pure gold, hybrid and bronze, taken from Unpaywall data (formerly OADOI) directly. Here, we improve on this by running all DOIs retrieved from WoS through Unpaywall data (using their web interface that allows batch checking of up to 10,000 DOIs at a time). Unlike WoS, Unpaywall data itself includes author-submitted versions in their green OA classification, resulting in more complete green OA levels. </p> <p>In addition, since our initial analysis of December 2017, Unpaywall data has considerably expanded its coverage of institutional repositories (IRs) (see <a href="https://unpaywall.org/sources">https://unpaywall.org/sources</a>). This now includes coverage of the IRs from all Dutch universities. </p> <p>Taken together, the current data show higher levels of green open access, including author-submitted versions, compared to our previous analysis. </p> <p>In this update, we include output (articles and reviews) from 2016 and 2017 for all 14 universities in the Netherlands. </p> <p>The following categories are distinguished (description taken from Piwowar at al., 2018, doi: <a href="https://doi.org/10.7717/peerj.4375">10.7717/peerj.4375</a>)</p> <ul> <li><strong>Pure gold</strong>: Published in an open-access journal (as defined by the DOAJ)</li> <li><strong>Hybrid</strong>: Free under an open license in a toll-access journal</li> <li><strong>Bronze</strong>: Free to read on the publisher page, but without a license</li> <li><strong>Green: </strong>Available from an institutional or disciplinary repository (including PubMedCentral)</li> </ul> <p>Data for Dutch universities were collected from Web of Science using the organization-enhanced field. Only articles and reviews were included. DOIs were extracted from the Web of Science export, run through the Unpaywall data <a href="https://unpaywall.org/products/simple-query-tool">Simple Query Tool</a>. From the resulting data from Unpaywall, OA classification was done using a simple formula in Excel (to be replaced by an R script in a future update). The Excel template used is included in this dataset, as is the OADOI API output for each Dutch university's article subset, and the lists of DOIs derived from Web of Science. The dataset also includes summarized data and three charts generated from these data, showing levels of different types of OA for 2016, 2017 and the two years compared. </p> <p>----------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p>
Data archive for Pepper, Bateson and Nettle, 'Telomeres as integrative markers of exposure to stress and adversity: A systematic review and meta-analysis'
<p>Data archive for the paper 'Telomeres as integrative markers of exposure to stress and adversity: A systematic review and meta-analysis' by Gillian Pepper, Melissa Bateson and Daniel Nettle. This version was uploaded in July 2018 after peer-review in the journal Royal Society Open Science. Compared to earlier version, it incorporates some minor error correction to the dataset, and reflects the revised analyses we performed after peer review. </p> <p>Our protocol and recording guide, which were preregistered on the Open Science Framework in 2016, are also included here, as is our PRISMA diagram.</p> <p>The data file 'unprocessed data' contains the data as extracted from the literature, with associations shown both as provided in the original papers, and converted to correlation coefficients. The algorithms for converting all the different associations to correlation coefficients are described in the flowchart and implemented in the R script 'effect conversion algorithms.r'.</p> <p>The data file 'processed data.csv' is the dataset analysed in the paper. Compared to 'unprocessed data.csv', it excludes: associations from studies of non-human animals; duplicate associations; a small number of associations from studies of medical treatments; and associations considered subparts or subscales of other associations. These exclusions are outlined in Methods section of the paper. In addition, in the processed data file, all correlations are aligned in direction so as to make them comparable (variable 'ValencedEffect'); and all associations are assigned to broad and fine categories.The script 'unprocessed to processed.r' makes the processed data file from the unprocessed one, or you can simply work from the processed one directly. </p> <p>The R script 'telomere metanalysis script RSOS REVISED.r' reproduces the analyses found in the paper.</p> <p>This version of the archive (July 17 2018) contains one small correction in the data files compared to all earlier versions. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.