Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
116
datasets available to search
ShareScore release 0.9.0
Dataset results
116 results for “Code Review”
Data and R code for Tansley review New Phytologist 2021: "An integrated framework of plant form and function: The belowground perspective"
<p>The files in this archive are related to the paper of Weigelt, Mommer, Andraczek et al. (2021) An integrated framework of plant form and function: The belowground perspective. Tansley Review New Phytologist. The paper developed and tested a new conceptual framework of plant form and function linking above and belowground traits of 2510 species. We found that an integrated, whole-plant trait space required as much as four axes. The two main axes represented the fast-slow ‘conservation’ gradient on which leaf and fine-root traits were well aligned, and the ‘collaboration’ gradient in roots. The two additional axes were separate, orthogonal plant size axes for height and rooting depth.</p> <p>This archives contains four files:</p> <ol> <li><strong>Weigelt et al.2021RCode.DataCleaning.txt</strong> - RCode for the complete data processing starting with the downloaded database files from the Plant Trait Database version 5.0 (TRY, Kattge et al. 2020), the Global Root Trait database (GRooT, Guerrero-Ramirez et al. 2020) and a small number of additional data files listed in Table S2 of the original paper. Additional information was later incorporated using FungalRoot Database (Soudzilovkaia et al. 2020), nodDB Database (Tedersoo et al. 2018) and a compiled dataset on rooting depth (Fan et al. 2017). The code processes, cleans and merges the data and produces a final table for PCA analysis of species specific mean traits. This final table is provided as a second file in this archive (Weigelt_et_al_2021_Main.PCA.Matrix.xlsx). A second part of the RCode.DataCleaning extracts species-specific individual trait data where root and shoot traits were measured on the same plant individual or plot. This data was compiled from 43 studies identified in Table S2 of the original publication. The final table for individual trait data is the third file in this archive (Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx).</li> <li><strong>Weigelt_et_al_2021_Main.PCA.Matrix.xlsx</strong> – Datafile with species-specific global mean trait data for 17 traits of 2510 species with at least one root and one shoot trait available. Meta-data is provided in the data file.</li> <li><strong>Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx</strong> – Datafile with species-specific trait data where root and shoot traits were measured on the same individual or plot for 6 traits of 455 species. Meta-data is provided in the data file.</li> <li><strong>Weigelt et al.2021RCode.Analysis.txt – </strong>RCode for all analyses and figures provided in the paper for both the species mean and individual based dataset. The Code is annotated to help reproducibility of the analysis.</li> </ol>
Datasets and codes for the peer review article "Human and natural impacts on the U.S. freshwater salinization and alkalinization: A machine learning approach"
<p>Ongoing salinization and alkalinization in U.S. rivers have been attributed to inputs of road salt and effects of human-accelerated weathering in previous studies. Salinization poses a severe threat to human and ecosystem health, while human derived alkalinization implies increasing uncertainty in the dynamics of terrestrial sequestration of atmospheric carbon dioxide. A mechanistic understanding of whether and how human activities accelerate weathering and contribute to the geochemical changes in U.S. rivers is lacking. To address this uncertainty, we compiled dissolved sodium (salinity proxy) and alkalinity values along with 32 watershed properties ranging from hydrology, climate, geomorphology, geology, soil chemistry, land use, and land cover for 226 river monitoring sites across the coterminous U.S. Using these data, we built two machine-learning models to predict monthly-aggregated sodium and alkalinity fluxes at these sites. The sodium-prediction model detected human activities (represented by population density and impervious surface area) as major contributors to the salinity of U.S. rivers. In contrast, the alkalinity-prediction model identified natural processes as predominantly contributing to variation in riverine alkalinity flux, including runoff, carbonate sediment or siliciclastic sediment, soil pH and soil moisture. Unlike prior studies, our analysis suggests that the alkalinization in U.S. rivers is largely governed by local climatic and hydrogeological conditions.</p>
SHAPE-ID Literature Review dataset: journal occurrences with ASJC codes
<p><strong>Background and methodology:</strong></p> <p>The dataset consists of a list of 2202 journal titles represented in the <a href="https://doi.org/10.5281/zenodo.4034507">SHAPE-ID Literature Review bibliography</a>, prepared for the purposes of quantitative analysis.</p> <p>The list of journals is based on 3955 journal articles in the bibliography dataset that had an International Standard Serial Number (ISSN). To each journal title the project team attributed:</p> <p>- a weight factor based on how many articles from the given journal featured in bibliography dataset</p> <p>- at least one <a href="https://service.elsevier.com/app/answers/detail/a_id/15181/supporthub/scopus/">All Science Journal Classification</a> (ASJC) code, representing different scientific disciplines</p> <p>- a country of publication. </p> <p>In case of 1853 of those journal titles, the attribution was automatised (we matched the ISSNs of journal titles in our sample against the Scopus Sources list from February 2019). In case of the remaining 349 titles the attribution was accomplished manually, based on the information available in SCOPUS, Web of Science, JSTOR, Information Matrix for the Analysis of Journals (MIAR) and ISSN databases.</p> <p><strong>Description of the file:</strong></p> <p>This is a csv file containing a list of 2202 journal titles represented in the SHAPE-ID Literature Review bibliography, with country of publication and ASJC codes assigned. </p> <p>The file is formatted as follows:</p> <p>Column A: ISSN of the journal</p> <p>Column B: information on how country and ASJC codes were attributed. Value “N” indicates automatic attribution based on match with Scopus list of sources. Other values indicate manual attribution. Values WOS, SCOPUS, JSTOR indicate source of information. Valu “Y” indicates that information was compiled based on multiple sources. </p> <p>Column C: numeric values correspond to the weight factor, i.e. number of time articles from each journal featured in the SHAP-ID Literature Review bibliography. </p> <p>Column D: SHAPE-ID Zotero bibliography identifier.</p> <p>Column E: Journal title</p> <p>Column F: The country of publication</p> <p>Columns G-AD: ASJC codes (numeric and word values) associated with journal entries. </p>
Understanding the Publish-Review-Curate (PRC) Model of Scholarly Communication - Data and Code
<p>Summary data for the number of articles submitted to publish-review-curate platforms as of August 2024 (Figure 1) [Update 14 Nov 2024: Added JMIRx. Data still from August 2024]</p> <p>Summary data for the number of articles reviewed by review platforms (Figure 2)</p> <p>Analysis code to produce Figures 1 and 2</p> <p>Code to extract articles for inclusion in data</p>
Reproduction package for paper "How far are we from reproducible research on code smell detection? A systematic literature review"
<p>Checklist and data extracted from publications analyzed for "How far are we from reproducible research on code smell detection? A systematic literature review" paper, together with processing scripts and calculations of Cohen's Kappa.</p> <p>Paper that describes details of the data is available here: https://doi.org/10.1016/j.infsof.2021.106783</p>
Data and Material for 'Less is More: Supporting Developers in Vulnerability Detection during Code Review'
<p>Data and Material supporting the paper 'Less is More: Supporting Developers in Vulnerability Detection during Code Review'.</p>
Privacy-by-Design Maturity Model: literature review, coding, model creation and evaluation
<p>Results from two multivocal literature reviews (MLRs) and subsequent coding, formulation of capabilities and dependencies, creation of maturity matrix and evaluation results. Used in the creation of a PbD domain model and extraction of core activities for PbD in Information Systems design. Part of the <a href="https://www.privacymaturity.org/" target="_blank" rel="noopener">Privacy-by-Design Maturity</a> research project by the <a href="https://www.uu.nl/en/research/ai-labs/ai-lab-for-the-public-services" target="_blank" rel="noopener">AI Lab for Public Services</a>.</p>
Data and material for: "Code review for newcomers: is it different?"
<p>Data and material for the paper "Code review for newcomers: is it different?", published in the Proceedings of the 11th International Workshop on Cooperative and Human Aspects of Software Engineering (CHASE 2018).</p>
The Upper Bound of Information Diffusion in Code Review
<p>More details on <a href="https://github.com/michaeldorner/information-diffusion-boundaries-in-code-review">https://github.com/michaeldorner/information-diffusion-boundaries-in-code-review</a></p>
Does Code Review Speed Matter for Practitioners?
<p>This dataset contains the following:</p> <ul> <li>The results from a survey about code velocity.</li> <li>The R script to analyze the survey data.</li> <li>The survey questionnaire.</li> </ul>
Mapping literature reviews on coral health: A review map, critical appraisal, and bibliometric analysis - Data and Code
<p>Data, code, and supplementary materials for "Mapping literature reviews on coral health: A review map, critical appraisal, and bibliometric analysis" by Burke et al., published in Ecological Solutions and Evidence.</p>
spanichella/RP_EMSE_MCR_2019 v.1.0.1 Second release of of the replication Package for the paper "An Empirical Investigation of Relevant Changes and Automation Needs in Modern Code Review".
<p>Replication Package for the paper "An Empirical Investigation of Relevant Changes and Automation Needs in Modern Code Review"</p> <p>Structure</p> <pre><code>project_raw_data/ gerrit_review_comments.csv gerrit_review_changes.csv survey_raw_data/ google_forms_survey.pdf google_forms_survey.csv RQ1_taxonomy_mcr/ RQ1_inception_phase/ initial_taxonomy.pdf intermediate_taxonomy.pdf RQ1_definition_phase/ Q1.2_evaluation_survey.csv cram_classified.csv cram.pdf RQ2_automation_needs/ Q2.1-Q2.5_evaluation_survey.xlsx Q2.6-Q2.7_evaluation_survey.xlsx Q2.1-Q2.7_question_index.csv Now "RQ3_automated_support/" contain the results concerning RQ2.1 in the paper. content explained in the README.md file located in "RQ3_automated_support/README.md" </code></pre> <p>Contents of the Replication Package</p> <p><strong>project-raw-data/</strong> contains the data used for the creation of our taxonomies, it includes information about the ten open-source projects.</p> <ul> <li><code>gerrit_review_comments.csv</code> - information about all in-line review comments used for this paper</li> <li><code>gerrit_review_changes.csv</code> - information about all patches analyzed that contain the in-line comments</li> </ul> <p><strong>survey_raw_data/</strong> contains information about the survey conducted for the paper.</p> <ul> <li><code>google_forms_survey.pdf</code> - the distributed <em>Google Forms</em> of our survey</li> <li><code>google_forms_survey.csv</code> - all survey answers obtained from 52 survey participants</li> </ul> <p><strong>RQ1_taxonomy_mcr/</strong> contains information and data about the elicited taxonomies in our paper (RQ1).</p> <ul> <li><strong>RQ1_inception_phase/</strong> <ul> <li><code>initial_taxonomy.pdf</code> - initial taxonomy obtained in the inception phase of our paper</li> <li><code>intermediate_taxonomy.pdf</code> - intermediate taxonomy after integrating and merging the initial taxonomy with the one by Beller <em>et al</em> [1]</li> </ul> </li> <li><strong>RQ1_definition_phase/</strong> <ul> <li><code>Q1.2_evaluation_survey.csv</code> - relevant survey feedback with additional taxonomy categories integrated into <em>CRAM</em></li> <li><code>cram_classified.csv</code> - classified review comments (<code>gerrit_review_comments.csv</code>) into <em>CRAM</em></li> <li><code>cram.pdf</code> - <em>CRAM</em> taxonomy</li> <li><code>cram_classified_with_frequency_information2020.xls</code> - it contain the information used to compute the frequency of CRAM changes, derived by the analysis of the 211 commits</li> </ul> </li> </ul> <p><strong>RQ2_automation_needs/</strong> contains the encoded evaluation of the survey question Q2.1-2.7 for RQ2</p> <ul> <li><code>Q2.1-Q2.5_evaluation_survey.xlsx</code> - the encoded evaluation of the survey questions Q2.1-Q2.5 (used for the <em>Automation Needs</em> Section in the paper) and contains the following sheets: <ul> <li><strong>all Findings</strong>: Very detailed findings matrix distilled from all answers concerning possible automated solutions or general possibilities to achieve automation in MCR. Every feedback was analyzed and decomposed into single findings. These findings are grouped, into categories of our Taxonomy of Code Changes in MCR (CRAM). Red represent in the feedback-text where the corresponding category was distilled from.</li> <li><strong>unique Findings</strong>: As one participant could mention the same categories/solutions in multiple feedbacks, the following matrix is cleaned of any duplication of participant answers. Multiple feedbacks containing the same information by one participant were removed, leaving only distinct occurances.</li> <li><strong>aggregated by Solution</strong>: Aggregated feeback clustered into abstracted solutions and the number of times participants mentioned the solution.</li> <li><strong>aggregated by Taxonomy</strong>: Aggregated feeback grouped by low-level categoried in CRAM.</li> </ul> </li> <li><code>Q2.6-Q2.7_evaluation_survey.xlsx</code> - the encoded evaluation of the survey questions Q2.6-Q2.7 (used for the <em>Automation Needs</em> Section in the paper) and contains the following sheets: <ul> <li><strong>all Findings</strong>: Very detailed findings matrix distilled from all answers concerning possible techniques, approaches and data to achieve automation in MCR. Every feedback was analyzed and decomposed into single findings. These findings are grouped, into categories of our Taxonomy of Code Changes in MCR (CRAM). Red represent in the feedback-text where the corresponding category was distilled from.</li> <li><strong>unique Findings</strong>: As one participant could mention the same categories/solutions in multiple feedbacks, the following matrix is cleaned of any duplication of participant answers. Multiple feedbacks containing the same information by one participant were removed, leaving only distinct occurances.</li> <li><strong>aggregated by low-level taxonomy</strong>: Aggregated mentionings of approaches/data by developers in the survey grouped by low-level taxonomy category.</li> <li><strong>aggregated by high-level taxonomy</strong>: Aggregated mentionings of approaches/data by developers in the survey grouped by high-level taxonomy category.</li> </ul> </li> <li><code>Q2.1-Q2.7_question_index.csv</code> - table of IDs given to each participant-question pair for Q2.1-Q2.7 in order to trace back the feeback.</li> <li><code>cram_survey-with_criticality_and_feasibility2020.xls</code> and <code>cram_survey-with_relevance_and_completeness_information2020.xls</code>: they contain we results of the survey, involving 14 additional participants (12 developers and 2 researchers), not involved in the aforementioned survey, and performed to qualitatively assess the relevance and completeness of the identified MCR change types as well as assess how critical and feasible to implement are some of the identified techniques to support MCR activities.</li> </ul> <p><strong>RQ2_1_automated_support/</strong> (or <strong>RQ3_automated_support/</strong> )- content explained in the README.md file located in "RP_EMSE_MCR_2019/tree/master/EMSE_MCR_2019/RQ3_automated_support/README.md"</p> <p>References</p> <p>[1] Moritz Beller, Alberto Bacchelli, Andy Zaidman, and Elmar Juergens. 2014. Modern code reviews in open-source projects: which problems do they fix?. In Proceedings of the 11th Working Conference on Mining Software Repositories (MSR 2014). ACM, New York, NY, USA, 202-211. DOI: <a href="http://dx.doi.org/10.1145/2597073.2597082">http://dx.doi.org/10.1145/2597073.2597082</a></p>
Information Needs in Contemporary Code Review - Appendix
<p>Contemporary code review is a widespread practice used by software engineers to maintain high software quality and share project knowledge. However, conducting proper code review takes time and developers often have limited time for review. In this paper, we aim at investigating the information that reviewers need to conduct a proper code review, to better understand this process and how research and tool support can make developers become more effective and efficient reviewers. Previous work has provided evidence that a successful code review process is one in which reviewers and authors actively participate and collaborate. In these cases, the threads of discussions that are saved by code review tools are a precious source of information that can be later exploited for research and practice. In this paper, we focus on this source of information as a way to gather reliable data on the aforementioned reviewers’ needs. We manually analyze 900 code review comments from three large open-source projects and organize them in categories by means of a card sort. Our results highlight the presence of seven high-level information needs, such as knowing the uses of methods and variables declared/modified in the code under review. Based on these results we suggest ways in which future code review tools can better support collaboration and the reviewing task. Appendix material.</p>
Code and data to support 'Street view imagery for built environment auditing: a systematic review'
<p>Code and data to support the manuscript entitled 'Street view imagery for built environment auditing: a systematic review'</p>
Coding frame and dataset for the study: Land-use governance: The interplay of social, market, and policy drivers – A global systematic review
<p>This file entails the coding frame and dataset used to conduct the systematic literature review "Land-use governance: The interplay of social, market, and policy drivers – A global systematic review".</p> <p>This study was first published as Chapter 2 of the PhD Dissertation "From soil to society - Rethinking governance for multifunctional land use and management" (Elsa L. Dingkuhn, 2025), and in a modified form in the journal Earth System Governance (Dingkuhn et al. 2025).</p> <p>The file consists of three sheets:<br>- Coding frame: Includes coding instructions and definitions used to extract and categorize the data.<br>- Variables: A list of dataset variables with explanations.<br>- List of included studies: The 81 studies from which the data was sourced.<br>- Data_List of observations: The dataset itself, consisting of 718 observations extracted from the included studies.</p>
Systematic review reveals sexually antagonistic knockouts in model organisms data and code
<p>R Code and data for manuscript titled "Systematic review reveals sexually antagonistic knockouts in model organisms".</p> <p>Drosophila data is from Ruzicka, F., Hill, M.S., Pennell, T.M., Flis, I., Ingleby, F.C., Mott, R., Fowler, K., Morrow, E.H., Reuter, M., 2019. Genome-wide sexually antagonistic variants reveal long-standing constraints on sexual dimorphism in fruit flies. PLoS Biol. 17, e3000244.</p> <p>https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3000244</p> <p>Human data is from Harper, J.A., Janicke, T., Morrow, E.H., 2021. Systematic review reveals multiple sexually antagonistic polymorphisms affecting human disease and complex traits. Evolution 75, 3087–3097. https://doi.org/10.1111/evo.14394</p> <p>https://onlinelibrary.wiley.com/doi/full/10.1111/evo.14394</p>
Dataset for Code Review Guidelines for GUI-based Testing Artifacts
<p>The Excel file contains meta-data about collected white and gray literature, applied inclusion/exclusion criteria, the code system, and a list of identified guidelines.</p>
R code and associated data for: A review of riverine ecosystem service quantification: research gaps and recommendations
<p>This publication contains the R code and associated data used in the Journal of Applied Ecology publication entitled "A review of riverine ecosystem service quantification: research gaps and recommendations". </p>
Dataset of smell comments in Code Review Discussions
<p>The raw data contains 104,321 code review comments, with records between January 2014 to August 2023. After a keyword search, 18,850 comments were manually analyzed by 26 developers. The analyzed data resulted in 3,798 smell comments. This meticulously curated dataset was used to collect 4,058 more smell comments through semantic search, comprising a total of 7,856 smell comments, which represents 13,27% of the original data collected. </p>
Program Comprehension Challenges in Software Code Review
<p>Software engineers spend more time understanding code than writing it (with up to 70% of their time being devoted to this). A key activity where developers spend a lot of time reading and understanding code is software code review. Yet, little is known on the challenges concerning program comprehension during software code review. This study provides insight into the types of comprehension challenges that occur during code reviews and their causes. We find that missing design rationale is the most common reason for comprehension challenges in software code review. Comprehension challenges occur most commonly around five topics: “program logic,” “code design,” “defensive coding,” “condition checking,” and “concurrency”. We also show that machine learning (ML) can be used to automatically detect comprehension challenges with 74.3% precision and a 66.7% recall. We discuss potential improvements to code review support tools based on our findings.</p> <p>These files contain the trained ML algorithms and all the data collected and used for this study.</p> <p> </p> <p>Data was collected in 2018 with analysis performed in 2018/2019, but completion was delayed due to impacts of COVID19.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.