Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
41,236
datasets available to search
ShareScore release 0.9.0
Dataset results
41,236 results for “Review of reviews”
Design Methodologies and Engineering Applications for Ecosystem Biomimicry: An Interdisciplinary Review Spanning Cyber, Physical, and Cyber-Physical Systems
<p>This is the data for the interdisciplinary review on ecosystem biomimicry for engineering applications. </p>
Incidence and Characteristics of Adverse Events in Paediatric Inpatient Care: a Systematic Review and Meta-Analysis
<p>This is the open data repository for the connected systematic review and meta-analysis.</p> <p>Data sets for the meta-analysis.</p> <p>Data collection file with all the information extracted from the included studies.</p> <p>QAT file with the information from the quality assessment tool (QAT) for all included studies.</p> <p>ReadMe with information on data sets and updates.</p> <p>Codebooks for data sets.</p> <p>R Code for the analysis</p>
IMDB Reviews
<p>IMDB Reviews: contains 348,415 user reviews about 50,000 movies. The scores for the movies, in a range [0,10], were discretized so that 10 classes are considered for classification. This is a highly imbalanced dataset.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition.</p>
Data from Time since liver transplantation and immunosuppression withdrawal outcomes: a systematic review with individual patient data meta-analysis
<p>This record provides one CSV file containing anonymized individual patient data (IPD) of pre-withdrawal times (in days) of liver transplant recipients that underwent immunosuppression (IS) withdrawal. Collection and publication of anonymized data was approved by the Ethics Committee Northwest and Central Switzerland. Patients of 15 primary studies are stratified by successfully reaching the state of IS-free operational tolerance (OT) or by developing signs of immunological rejection (non-OT).</p>
"Determinants of football players' valuation: a systematic review" datasets
<p>3 interdependent tables are available and are the materials used for a systematic review of the determinants of football players' valuation. </p> <p>- model_specifications is a table where each row represents one of the 111 model specifications included in the systematic review. Characteristics of the article from which specification was retrieved (title, year, authors, journal), attributes of the specification (sample size, population, econometric modeling, etc.), the significance levels of included variables, and the associated coefficients for significant variables. </p> <p>- model_specifications_dictionnary precise the names of the columns of the table model_specifications.</p> <p>- variables_definitions_and_classification is a table that details all the 471 variables used in the 29 articles analysed with definitions quoted from articles when possible and presents a classification of these variables into 6 categories and several subcategories. </p>
A Systematic Literature Review of Machine Learning for Uncovering Software Faults and Failures
<p>This data set contains the results of an extensive, systematic literature review on the use of machine learning (ML) for uncovering software faults and failures. Covering the period of 2019 to 2022, this literature review identifies 874 relevant publications, classified into six distinct quality assurance tasks. Results show a compound annual growth rate (CAGR) of relevant publications of 38% over the last five years.</p> <p>This literature review particularly analyzed in how far these relevant papers leverage synergies between different quality assurance tasks. Results show that only 3% of all relevant papers leverage such synergies, indicating ample opportunities for future research. For example, a single type of quality assurance activity may not suffice to deliver the expected software quality. Ideally, one would use a suitable combination of different types of activities – such as combining dynamic testing with static code analysis. Also, leveraging the synergies between different quality assurance activities can increase the effectiveness of the individual activities. For example, having a good estimate of the fault density of a software component (e.g., using deep learning-driven fault prediction techniques) could help optimize and prioritize testing effort and budget.</p>
Literature Datasets for the publication "Systematic Review: Prevalence and Practices of Immunofluorescent Cell Image Processing"
<p>This dataset contains the CSV files returned from PubMed searches used to complete a Systematic Review of Image Processing Publication Practices for methods applied to immunofluorescent images of all CNS cells. <br> <br> The file names are organized "date_supplementarytablenumber" followed by the appropriate search terms. </p>
Research data for Critical review of heat pump prototype operation and required modications
<p>In this dataset, the data of the first experimental test campaign of the CO2-ice heat pump is shared. The report, which analyzes the data and makes the necessary explanations, has already been shared as a "Critical review of heat pump prototype operation and required modifications (Deliverable: D5.5)".</p>
A Systematic Review on the Visualization of Avatars and Agents in AR & VR displayed using Head-Mounted Displays
<p>This repository contains the data for the article "Systematic Review on the Visualization of Avatars and Agents in<br> AR & VR displayed using Head-Mounted Displays" including a BibTex file and supporting figures.</p>
Metrics and peer review agreement at the institutional level - Data
<p>This data is released to accompany the paper:</p> <p>Traag, VA, Malgarini, M and Sarlo, S (2020) Metrics and peer review agreement at the institutional level.</p>
Replication package for Decomposition of Monolithic Applications into Microservices Architectures: A Systematic Review
<p><strong>Replication Package</strong></p> <p><strong>Title:</strong></p> <p>Replication package for Decomposition of Monolithic Applications into Microservices Architectures: A Systematic Review.</p> <p><strong>Authors:</strong></p> <p>Yalemisew Abgaz, Andrew McCarren, Peter Elger, David Solan, Neil Lapuz, Marin Bivol, Glenn Jackson, Murat Yilmaz, Jim Buckley, and Paul Clarke</p> <p><strong>Year</strong></p> <p>This replication package was initially generated in 2022 and following feedback from reviewers, it is revised in 2023.</p> <p>This package contains two files and three folders that provide additional insight for researchers who wish to replicate our work or who would like to expand the review in the future.</p> <p><strong>Files</strong></p> <ul> <li>The file contains a detailed description of the literature search outlining the steps and the results obtained.</li> <li>Readme.md: A readme file (this file).</li> </ul> <p><strong>Folders</strong></p> <ul> <li>A folder containing the search results and the refinement steps. It contains four files listing studies included in the refinement process (Refinement_Step_1 to Refinement_Step_4) and a master file combining all the steps in one file. The master file contains detailed information about how the refinement steps are executed and all the intermediate results following each refinement step. Users may explore by expanding the filters in Refinement Step 2 (J), Refinement Step 3 (M) and Refinement Step 4 § columns in the master sheet. Readers can also directly go to the sheets that contain the selected studies in any of the refinement stages. A description of each file is also included in the Literature_Search_Strategy.pdf file.</li> <li>This folder contains the list of studies included in the snowballing process, including the last two refinement steps (Refinement_Step_5 and Refinement_Step_6). The snowballing master sheet contains studies extracted using the snowballing process and the data cleaning and filtering criteria used. The different sheets also contain the selected studies at each stage of the snowballing process.</li> </ul> <ul> <li>This folder contains the data extracted from the selected literature by employing the Systematic Review and Ground Theory. The two files included in this folder contain the data extraction template and the data extracted from the 35 selected studies including some intermediate notes.</li> </ul> <p>If you have further questions regarding the survey, feel free to contact us via e-mail. <a href="mailto:Yalemisewm.abgaz@dcu.ie">Yalemisewm.abgaz@dcu.ie</a>.</p> <p> </p>
TripAdvisor Vietnam Hotel Reviews
<p>The TripAdvisor Vietnam Hotel Reviews Dataset is a comprehensive collection of user-generated reviews from the popular online travel platform TripAdvisor. This dataset offers valuable insights into the experiences, opinions, and ratings provided by individuals who have stayed at various hotels across Vietnam.</p> <p>The dataset encompasses many hotels in different cities and regions of Vietnam, including popular tourist destinations such as Hanoi, Ho Chi Minh City, Da Nang, Nha Trang, and more. The reviews cover a diverse spectrum of accommodation types, ranging from budget guesthouses to luxurious resorts, providing a comprehensive representation of the Vietnamese hospitality industry.</p> <p>Each review entry in the dataset includes a rich set of information, offering researchers, developers, and data analysts an in-depth understanding of hotel performance and customer satisfaction. Key attributes of the dataset include:</p> <ol> <li> <p>Review Text: The actual text of the review left by the user, which contains detailed descriptions, opinions, and feedback about their hotel experience.</p> </li> <li> <p>Rating: The overall rating provided by the reviewer, typically ranging from 1 to 5 stars, reflects their satisfaction level with the hotel.</p> </li> <li> <p>Date: The review was posted, enabling temporal analysis and tracking changes over time.</p> </li> <li> <p>Location: The hotel's geographic location allows researchers to analyze regional variations in hotel performance and customer preferences.</p> </li> </ol> <p>The TripAdvisor Vietnam Hotel Reviews Dataset is valuable for various applications, including sentiment analysis, opinion mining, natural language processing, customer behavior analysis, recommender systems, and more. Researchers can leverage this dataset to gain deep insights into customer experiences, identify patterns, trends, and sentiments, and develop data-driven strategies for the Vietnamese hotel industry.</p>
Project "Public services management system to improve the quality and accessibility of services" (01.2.2-LMT-K-718-03-0019) literature review screening results
<p>The results of the keyword query in Scopus search with the abstracts were screened using <i>abstractr </i>platform at <a href="http://abstrackr.cebm.brown.edu">http://abstrackr.cebm.brown.edu</a>. Four reviewers reviewed intersecting subsets of the overall list of publications in separate reviews, therefore duplicate records in the file are possible. The results from four reviews were combined into one file using functionality of <i>abstractr </i>platform. The reviews were finalized in February, 2022. Majority of the publications from the Scopus query results were automatically assigned low relevance scores thanks to the active learning algorithm used by <i>abstractr</i> and therefore were not reviewed manually.</p><p>Notes on the columns of the dataset:</p><ul><li>(internal) id - internal id added by <i>abstractr.</i></li><li>(source) id - Scopus document id followed by underscore and '1' (if publication has DOI) or 'n' (if publication has no DOI).</li><li>keywords - authors' keywords and Scopus keywords concatenated from the Scopus query results.</li><li>abstract - an abstract of a publication from the Scopus query results.</li><li>title - title of a publication from the Scopus query results.</li><li>journal - journal of a publication from the Scopus query results.</li><li>authors - authors of a publlication from the Scopus query results.</li><li>consensus - for publications reviewed by multiple reviewers the consensus decision is signified by '1', no consensus - by 'x' and unable to asses consensus by 'o'. These codes are generated by <i>abstrackr.</i></li><li>eg - the first reviewer id. Code '1' means the publication was selected based on title and abstract, '0' - unsure, '-1' rejected.</li><li>dj - the second reviewer id. Code '1' means the publication was selected based on title and abstract, '0' - unsure, '-1' rejected.</li><li>rp - the third reviewer id. Code '1' means the publication was selected based on title and abstract, '0' - unsure, '-1' rejected.</li><li>mp - the fourth reviewer id. Code '1' means the publication was selected based on title and abstract, '0' - unsure, '-1' rejected.</li><li>count (+1) - integer, the number of reviewers selecting the publication for fulltext reading. Calculated from 'eg ','dj ','rp' and 'mp' columns.</li><li>count (0) - integer, the number of reviewers not sure of selecting the publication for fulltext reading. Calculated from 'eg ','dj ','rp' and 'mp' columns.</li><li>count (-1) - integer, the number of reviewers not selecting the publication for fulltext reading. Calculated from 'eg ','dj ','rp' and 'mp' columns.</li><li>at leat once selected - binary integer, representing the final decision rule to select publications for fulltext reading and further analysis.</li></ul><p>The data were collected as a part of the project "Public services management system to improve the quality and accessibility of services" ("Viešųjų paslaugų vadybos sistema paslaugų kokybei ir prieinamumui gerinti"), grant no. 01.2.2-LMT-K-718-03-0019, funded by the Lithuanian research council.</p>
The Upper Bound of Information Diffusion in Code Review
<p>More details on <a href="https://github.com/michaeldorner/information-diffusion-boundaries-in-code-review">https://github.com/michaeldorner/information-diffusion-boundaries-in-code-review</a></p>
Dataset: A Systematic Literature Review on the topic of High-value datasets
<p>This dataset contains data collected during a study ("T<a href="http://https://arxiv.org/abs/2305.10234">owards High-Value Datasets determination for data-driven development: a systematic literature review</a>") conducted by Anastasija Nikiforova (University of Tartu), Nina Rizun, Magdalena Ciesielska (Gdańsk University of Technology), Charalampos Alexopoulos (University of the Aegean)<sup> </sup>and Andrea Miletič (University of Zagreb)<br> It being made public both to act as supplementary data for "Towards High-Value Datasets determination for data-driven development: a systematic literature review" paper (pre-print is available in Open Access here -><a href="http://https://arxiv.org/abs/2305.10234"> https://arxiv.org/abs/2305.10234</a>) and in order for other researchers to use these data in their own work.</p> <p><br> The protocol is intended for the Systematic Literature review on the topic of High-value Datasets with the aim to gather information on how the topic of High-value datasets (HVD) and their determination has been reflected in the literature over the years and what has been found by these studies to date, incl. the indicators used in them, involved stakeholders, data-related aspects, and frameworks. The data in this dataset were collected in the result of the SLR over Scopus, Web of Science, and Digital Government Research library (DGRL) in 2023.</p> <p> </p> <p>***Methodology***</p> <p>To understand how HVD determination has been reflected in the literature over the years and what has been found by these studies to date, all relevant literature covering this topic has been studied. To this end, the SLR was carried out to by searching digital libraries covered by Scopus, Web of Science (WoS), Digital Government Research library (DGRL).</p> <p>These databases were queried for keywords <em>("open data" OR "open government data") AND ("high-value data*" OR "high value data*"</em>), which were applied to the article title, keywords, and abstract to limit the number of papers to those, where these objects were primary research objects rather than mentioned in the body, e.g., as a future work. After deduplication, 11 articles were found unique and were further checked for relevance. As a result, a total of 9 articles were further examined. Each study was independently examined by at least two authors.</p> <p>To attain the objective of our study, we developed the protocol, where the information on each selected study was collected in four categories: (1) descriptive information, (2) approach- and research design- related information, (3) quality-related information, (4) HVD determination-related information.</p> <p> </p> <p>***Test procedure***<br> Each study was independently examined by at least two authors, where after the in-depth examination of the full-text of the article, the structured protocol has been filled for each study.<br> The structure of the survey is available in the supplementary file available (see Protocol_HVD_SLR.odt, Protocol_HVD_SLR.docx)<br> The data collected for each study by two researchers were then synthesized in one final version by the third researcher.</p> <p>***Description of the data in this data set***</p> <p>Protocol_HVD_SLR provides the structure of the protocol<br> Spreadsheets #1 provides the filled protocol for relevant studies.<br> Spreadsheet#2 provides the list of results after the search over three indexing databases, i.e. before filtering out irrelevant studies</p> <p>The information on each selected study was collected in four categories:<br> (1) descriptive information,<br> (2) approach- and research design- related information,<br> (3) quality-related information,<br> (4) HVD determination-related information</p> <p>Descriptive information <br> 1) Article number - a study number, corresponding to the study number assigned in an Excel worksheet<br> 2) Complete reference - the complete source information to refer to the study<br> 3) Year of publication - the year in which the study was published<br> 4) Journal article / conference paper / book chapter - the type of the paper -{journal article, conference paper, book chapter}<br> 5) DOI / Website- a link to the website where the study can be found<br> 6) Number of citations - the number of citations of the article in Google Scholar, Scopus, Web of Science<br> 7) Availability in OA - availability of an article in the Open Access<br> 8) Keywords - keywords of the paper as indicated by the authors<br> 9) Relevance for this study - what is the relevance level of the article for this study? {high / medium / low}</p> <p>Approach- and research design-related information<br> 10) Objective / RQ - the research objective / aim, established research questions<br> 11) Research method (including unit of analysis) - the methods used to collect data, including the unit of analy-sis (country, organisation, specific unit that has been ana-lysed, e.g., the number of use-cases, scope of the SLR etc.)<br> 12) Contributions - the contributions of the study<br> 13) Method - whether the study uses a qualitative, quantitative, or mixed methods approach?<br> 14) Availability of the underlying research data- whether there is a reference to the publicly available underly-ing research data e.g., transcriptions of interviews, collected data, or explanation why these data are not shared?<br> 15) Period under investigation - period (or moment) in which the study was conducted<br> 16) Use of theory / theoretical concepts / approaches - does the study mention any theory / theoretical concepts / approaches? If any theory is mentioned, how is theory used in the study?</p> <p>Quality- and relevance- related information <br> 17) Quality concerns - whether there are any quality concerns (e.g., limited infor-mation about the research methods used)?<br> 18) Primary research object - is the HVD a primary research object in the study? (primary - the paper is focused around the HVD determination, sec-ondary - mentioned but not studied (e.g., as part of discus-sion, future work etc.))</p> <p>HVD determination-related information <br> 19) HVD definition and type of value - how is the HVD defined in the article and / or any other equivalent term?<br> 20) HVD indicators - what are the indicators to identify HVD? How were they identified? (components & relationships, “input -> output")<br> 21) A framework for HVD determination - is there a framework presented for HVD identification? What components does it consist of and what are the rela-tionships between these components? (detailed description)<br> 22) Stakeholders and their roles - what stakeholders or actors does HVD determination in-volve? What are their roles?<br> 23) Data - what data do HVD cover?<br> 24) Level (if relevant) - what is the level of the HVD determination covered in the article? (e.g., city, regional, national, international)</p> <p><br> ***Format of the file***<br> .xls, .csv (for the first spreadsheet only), .odt, .docx</p> <p>***Licenses or restrictions***<br> CC-BY</p> <p> </p> <p>For more info, see README.txt</p>
RSE-AUNZ community review: responses to steering committee survey
<p>These are the anonymous responses to the steering committee survey conducted for a review of the Research Software Engineers Australia and New Zealand (RSE-AUNZ) community. <a href="https://doi.org/10.5281/zenodo.8098001">Read the report</a>.<br> <br> This survey is one of 2 surveys conducted for the review, the other being a community survey. <a href="https://doi.org/10.5281/zenodo.8098070">Read the results</a>.</p>
Dataset of "Early triadic interactions in the first year of life: A systematic review on object-mediated shared encounters"
<pre>Two files are available as a result of data extraction from the 51 studies included in the systematic review entitled: "Early triadic interactions in the first year of life: A systematic review on object-mediated shared encounters": (1) DataExtraction.csv -> Dataset resulting from data extraction. (2) DictionaryVariables.csv -> dictionary of the variables included in the dataset: Authors, doi, Year of publication, Study type, Design type, Data analysis strategy, Type of task, Objects, Infants' age (months), Context of interaction, and Country. </pre>
Rapid Review Dataset for Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team
<p>A collection of evidence briefings produced through a rapid literature review protocol the Department of Software Engineering and Research at Sandia National Laboratories. These briefings are described in our research paper, "Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team", which was accepted for publication at the 1st Annual Conference of the United States Research Software Engineer Association (US-RSE'23).</p> <p>Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525. SAND2023-06549O.</p>
Does Code Review Speed Matter for Practitioners?
<p>This dataset contains the following:</p> <ul> <li>The results from a survey about code velocity.</li> <li>The R script to analyze the survey data.</li> <li>The survey questionnaire.</li> </ul>
InTheMED WP2 Data Archive - Groundwater Quality Literature Review
<p>The data archive InTheMED_WP2_DS_GWQualityLitReview is part of Task 2.2 “Review and collect the available groundwater quantity and quality data sets in the MED region” and contains a literature review of groundwater quality data collected from various sites in Mediterranean countries. The data includes measurements of different water quality parameters, providing valuable insights into the characteristics of groundwater in different regions. The dataset was compiled from a literature review of published research articles.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.