Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
116
datasets available to search
ShareScore release 0.9.0
Dataset results
116 results for “code review”
Replication Package for the Paper: "Code Smells Detection via Code Review: An Empirical Study"
<p>This repository contains the data and results from the paper "Code Smells Detection via Code Review: An Empirical Study" submitted to ESEM 2020.</p> <p> </p> <p><strong>1. data folder</strong></p> <p>The data folder contains the retrieved 269 reviews that discuss code smells. Each review includes four parts: Code Change URL, Code Smell Term, Code Smell Discussion, and Source Code URL.</p> <p> </p> <p><strong>2. scripts floder</strong></p> <p>The scripts folder contains the Python script that was used to search for code smell terms and the list of code smell terms.</p> <ul> <li><em>smell-term/general_smell_terms.txt</em> contains general code smell terms, such as "code smell".</li> <li><em>smell-term/specific_smell_terms.txt</em> contains specific code smell terms, such as "dead code".</li> <li><em>smell-term/misspelling_terms_of_smell.txt</em> contains the misspelling terms of 'smell', such as "ssell".</li> <li><em>get_changes.py</em> is used for getting code changes from OpenStack.</li> <li><em>get_comments.py</em> is used for getting review comments for each code change.</li> <li><em>smell_search.py</em> is used for searching review comments that contain code smell terms.</li> </ul> <p> </p> <p><strong>3. project folder</strong></p> <p>The project folder contains the MAXQDA project files. The files can be opened by MAXQDA 12 or higher versions, which are available at https://www.maxqda.com/ for download. You may also use the free 14-day trial version of MAXQDA 2018, which is available at https://www.maxqda.com/trial for download.</p> <ul> <li><em>Data Labeling & Encoding for RQ2.mx12</em> is the results of data labeling and encoding for RQ2, which were analyzed by the MAXQDA tool.</li> <li><em>Data Labeling & Encoding for RQ3.mx12</em> is the results of data labeling and encoding for RQ3, which were analyzed by the MAXQDA tool.</li> </ul>
EEG datasets for healthcare: a scoping review - Data and Code
<p>This repository contains the data extracted for the scoping review "EEG datasets for healthcare: a scoping review" and the code used in the analysis.</p><p> </p>
Data and code for peer review - MEE-24-11-799
Open the record for dataset details and reuse information.
On the acceptance by code reviewers of candidate security patches suggested by Automated Program Repair tools - Dataset
<p>Dataset of the empirical experiment presented in the paper On the acceptance by code reviewers of candidate security patches suggested by Automated Program Repair tools. The dataset includes the participants' responses regarding their background, and the responses of the tasks from the experiment. </p>
Effective Teaching through Code Reviews: Patterns and Anti-Patterns
Open the record for dataset details and reuse information.
Accessibility of Low-Code Approaches: a Systematic Literature Review
Open the record for dataset details and reuse information.
Table-Matrix, containing all the codes generated from the inductive analysis process of the 78 articles that made up the final sample. (Not only Opportunity, but also Uncertainty: A systematic review of how entrepreneurship literature appropriates both constructs.)
Open the record for dataset details and reuse information.
Sistematic review on the code smell effect (2000-2017)
<p>Dataset of the paper "A systematic review on the code smell effect".</p>
Challenges in analysis of code review comments. Various BERTopic parameters and impact on coherence.
Open the record for dataset details and reuse information.
Advancing Automated Code Review Comment Generation Using Large Language Models
Open the record for dataset details and reuse information.
Towards Automating Code Review Activities - JSS Happy Hour Video
<p>Towards Automating Code Review Activities - JSS Happy Hour Video</p>
Data and Material for the master's thesis: Cost of code review goals and code review strategies
<p>Data and Material for the master's thesis: Cost of review goals and review strategies</p>
Fire code review sketch dataset
<p>Automatic assessments of building plans are uncommon in the early design stages, especially when schematic sketches are in raster format. Existing design evaluation tools, such as fire code reviewers, which are typically used in the late design stage, primarily evaluate vector format images that contain complete building information. These tools use conditional shape-embedding techniques to analyze the vector images. However, there are limitations to identifying and evaluating drawings through vector-shape relationships. Our research aimed to develop tools that can automatically assess schematic sketches in raster format to overcome the limitations of existing tools. We integrated a conditional shape-embedding tool, named Shape Machine, to assess vector images, with machine learning techniques, namely a Generative Adversarial Network (GAN), to assess raster sketches. This integration enables the evaluation of fire evacuation sketches in the early stages of the design process, thereby improving design efficiency and reducing costs. Moreover, in the future, this integration could allow the evaluation of designs in multiple image formats.</p>
Help Me to Understand this Commit! - A Vision for Contextualized Code Reviews
<p>Literature review results: snowball sampling results and categorization of papers.</p>
Integrating Visual Aids to Enhance the Code Reviewer Selection Process Replication Package
<p>Modern Code Review (MCR) is an integral part of a software development strategy that accelerates product quality by identifying defects, code smells, and other harmful practices. However, assigning appropriate reviewers to evaluate changed code during the review process remains challenging. While automated tools for reviewer assignments have limited impact in practice, the process often relies on manual investigation of project histories to retrieve knowledge of team members and their activities. Therefore, in this study, we present an approach to automatically assemble developers' information and visualize it meaningfully, which helps to choose appropriate reviewers. First, we propose three metrics that measure developers' collaboration, reviewers' expertise, and reviewers' workload and visualize them through networks. Second, we perform a case study of three popular open-source projects, where we compute and visualize each developers' information according to the proposed metrics. Finally, we conducted two online surveys to assess the developers' perceptions of the proposed visual benefits. The results show that the proposed method can assist in identifying relevant reviewers and be immensely helpful to new developers. Additionally, survey respondents expressed reliance on the efficacy of the visual aids in workload balancing and reducing review time. </p>
Code review regression analysis of open source GitHub projects
Open the record for dataset details and reuse information.
Dataset of "Primers or Reminders? The Effects of Existing Review Comments on Code Review"
<p>Dataset of "Primers or Reminders? The Effects of Existing Review Comments on Code Review".</p> <p>See README.md for more information. </p>
Dataset of the paper "Information Needs in Contemporary Code Review"
<p>Dataset of the paper "Information Needs in Contemporary Code Review", Proceedings of the ACM on Human-Computer Interaction 2, CSCW, Article 135 (November 2018).</p> <p>Read README.md for information on the content.</p>
Model code and data for "Mitigation of the double ITCZ syndrome in BCC-CSM2-MR through improving parameterizations of boundary-layer turbulence and shallow convection" by Lu et al., submitted to Geoscientific Model Development, https://doi.org/10.5194/gmd-2020-40, in review, 2020.
<p>Description of the files:</p> <p>“BCC_CSM2_MR.code.tar” contains the codes and run scripts for the medium-resolution Beijing Climate Center Climate System Model version 2 (BCC-CSM2-MR). Detailed description of the model refers to the paper “The Beijing Climate Center Climate System Model (BCC-CSM): the main progress from CMIP5 to CMIP6” by Wu et al., Geosci. Model Dev., 12, 1573–1600, https://doi.org/10.5194/gmd-12-1573-2019, 2019.</p> <p>“BCC_CSM2_MR.inputdata.tar” contains the input data needed to run the model.</p> <p>“REF_amip.rar” contains the output data from the REF_amip experiment.</p> <p>“NEW_amip.rar” contains the output data from the NEW_amip experiment.</p> <p>“REF_cmip.rar” contains the output data from the REF_cmip experiment.</p> <p>“NEW_cmip.rar” contains the output data from the NEW_cmip experiment.</p> <p>“UWMT_amip.rar” contains the output data from the UWMT_amip experiment.</p> <p>“mHack_amip.rar” contains the output data from the mHack_amip experiment.</p>
Identifying prevalent quality issues in code changes by analyzing reviewers' feedback
<p>The dataset contains python scripts (for data extraction, preprocessing) and dataset for the paper "Identifying prevalent quality issues in code changes by analyzing reviewers' feedback" Submitted to ENASE 2024.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.