Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
184
datasets available to search
ShareScore release 0.9.0
Dataset results
184 results for “Replication Study”
Opportunities and Security Risks of Technical Leverage: A Replication Study on the NPM Ecosystem
<p>Metadata, technical leverage, and vulnerability reports for 14,042 stable releases and scripts used in "Opportunities and Security Risks of Technical Leverage: A Replication Study on the NPM Ecosystem" paper.</p> <p> </p> <p> </p>
Replication package: 40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study
<p>Replication package | 40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study</p>
Exploring Gender Bias in Remote Pair Programming among Software Engineering Students: The Twincode Original Study and First External Replication (datasets)
<p>This repository contains the datasets of the original experiment (University of Seville, December 2021) and its first external replication (University of California, Berkeley, May 2022) of the Twincode exploratory study on the effects of gender bias in remote pair programming among software engineering students.</p>
Replication Kit for Paper: "Are There Any Unit Tests? An Empirical Study on Unit Testing in Open Source Python Projects"
<p>Replication Kit for the Paper "Are there any Unit Tests? An Empirical Study on Open Source Python Projects" by Fabian Trautsch, Jens Grabowski.</p> <p>You can cite the paper via:</p> <p>@inproceedings{trautsch2017there,<br> title={Are There Any Unit Tests? An Empirical Study on Unit Testing in Open Source Python Projects},<br> author={Trautsch, Fabian and Grabowski, Jens},<br> booktitle={Proceedings of the IEEE International Conference on Software Testing, Verification and Validation (ICST)},<br> pages={207--218},<br> year={2017},<br> organization={IEEE}<br> }</p> <p> </p> <p>Contents:<br> 1) Used version of the vcsSHARK<br> - located in “vcsSHARK”<br> 2) Used version of the testImpSHARK<br> - located in “testImpSHARK”<br> 3) Analysis implementations<br> - located in “testImpSHARK/testimpshark/analysis”<br> 4) Raw Data CSV Files<br> - located in “testImpSHARK/testimpshark/analysis/data<br> 5) Raw MongoDB<br> - located in “mongo_backup”</p> <p><br> Usage:<br> 1) Usage instructions for the vcsSHARK is given on its github homepage (http://ftrautsch.github.io/vcsSHARK/index.html) or directly in the “vcsSHARK/pyvcsshark/main.py” file</p> <p>2) Usage instructions for the testImpSHARK:<br> - if only one revision should be analyzed use “testImpSHARK/main.py”<br> - if all revisions should be analyzed use “testImpSHARK/execution.py”<br> - in both files concrete instructions can be found</p> <p>3) Each analysis file is commented. For some of them (rq1_boxplot.py and rq4.py the connection to the MongoDB must be changed). For the R files, the path to the data must be adapted. Otherwise, the files can be directly executed.</p> <p>4) The MongoDB can be restored via:<br> mongorestore --gzip --archive=smartshark040816.gz --db smartshark --host <HOST> --port <PORT> --username <USERNAME> --password <PASSWORD> --authenticationDatabase <AUTHENTICATION_DATABASE></p> <p><br> Tests:<br> 1) The tests can be run directly via the unittest framework of python: e.g., python -m unittest tests/test_common.py</p>
Evolution repeats itself in replicate long-term studies in the wild
<p>The extent to which evolution is repeatable remains debated. Here we study changes over time in the frequency of cryptic color-pattern morphs in 10 replicate long-term field studies of a stick-insect, each spanning at least a decade (across 30 years of total data). We find predictable 'up-and-down' fluctuations in stripe frequency in all populations, representing repeatable evolutionary dynamics based on standing genetic variation. A field experiment demonstrates that these fluctuations involve negative frequency-dependent natural selection (NFDS). These fluctuations rely on demographic and selective variability that pushes populations away from equilibrium, such that they can reliably move back towards it via NFDS. Finally, we show that the origin of new cryptic forms is associated with multiple structural genomic variants such that which mutations arise affects evolution at larger temporal scales. Thus, evolution from existing variation is predictable and repeatable, but mutation adds complexity even for traits evolving deterministically under natural selection.</p>
Supplementary materials to the article "Practices and attitudes toward replication in empirical translation and interpreting studies"
<p><strong>Abstract of the paper:</strong></p> <p>This article presents the results of three studies on practices in and attitudes toward replication in empirical translation and interpreting studies. The first study reports on a survey in which 52 researchers in translation and interpreting with experience in empirical research answered questions about their practices in and attitudes toward replication. The survey data were complemented by a bibliometric study of publications indexed in the Bibliography of Interpreting and Translation (BITRA) (Franco Aixelá 2001–2019) that explicitly stated in the title or abstract that they were derived from a replication. In a second bibliometric study, a conceptual replication of Yeung’s (2017) study on the acceptance of replications in neuroscience journals was conducted by analyzing 131 translation and interpreting journals. The article aims to provide evidence-based arguments for initiating a debate about the need for replication in empirical translation and interpreting studies and its implications for the development of the discipline.</p> <p><strong>Description of the supplementary materials:</strong></p> <ul> <li>Material 1: questionnaire in English</li> <li>Material 2: questionnaire in Spanish</li> <li>Material 3: results of both questionnaires</li> <li>Material 4: corpus of identified replications in Translation and Interpreting Studies (TIS)</li> <li>Material 5: acceptance of replications in TIS journals</li> </ul>
Replication package for "An Exploratory Study on Default Class Openness in Java and Kotlin"
<p>This dataset includes scripts, text files, and cached CSV/Parquet or raw TXT data files used to generate all analysis and results from the paper. A <strong>README.md</strong> file is included in <strong>replication-pkg.zip</strong> for details on using the scripts.</p> <p>If you only want to inspect the figures, you do not need a data ZIP.</p> <p>If you want to simply re-generate the figures without changes, download <strong>data-cached.zip</strong>. If you want to make any sort of change to the analyses, you will want to download <strong>data-csv.zip</strong> and/or <strong>data-raw.zip</strong>. The CSV files are converted versions of the raw TXT files.</p>
Replication package for "An Empirical Study of Q&A Websites for Game Developers"
<p><strong>Replication package for the paper "An Empirical Study of Q&A Websites for Game Developers"</strong></p> <p>This repository contains the datasets and scripts used to replicate the results from the paper "An Empirical Study of Q&A Websites for Game Developers".</p> <p>This is an exact copy of the repository on GitHub: https://github.com/asgaardlab/done-21-arthur-gamedev_qa_websites-code</p> <p><strong>Replication data</strong></p> <p>The datasets used to replicate the results for the paper can be found in the data directory (data/). These are the datasets we obtained after running all of the notebooks in this repository.</p> <p>Two of the studied websites are owned by companies (Epic and Unity) and we are not legally allowed to share the textual contents of the questions and answers as they are considered intellectual property. Therefore, instead of sharing the content of those posts, we included the URLs to all of the pages where the information used in the paper can be found, so that they can be crawled by future researchers.</p> <p>This is not an issue for Stack Overflow and the Game Development Stack Exchange, since that data is provided by Stack Exchange in the Stack Exchange Data Dump (https://archive.org/details/stackexchange).</p> <p><strong>Survey data:</strong> Unfortunately, our University's ethics board only allows us to share the survey responses in aggregated format, which is done in the paper. In this repository, we added the list of communities in which we shared the survey (data/surveyed_communities.csv).</p> <p><strong>Using this repository</strong></p> <p>If you are using the datasets provided in this repository, you just need to run the analysis notebook (code/analysis/paper_results.ipynb) to obtain the results as shown in the paper.</p> <p>Otherwise, if you want to run the whole pipeline from scratch, follow these steps:</p> <p>1. Download the data from Unity Answers and the UE4 AnswerHub from their websites (you can use the URLs provided in our datasets). Parse the HTML pages and extract the required information.</p> <p>2. Download the data from Stack Overflow and the Game Development Stack Exchange from the Stack Exchange Data Dump (https://archive.org/details/stackexchange). Run the notebooks to process the XML files from the Stack Exchange data dump (code/process_xml). For Stack Overflow, run the select_gamedev_posts.ipynb (code/process_xml/stackoverflow/select_gamedev_posts.ipynb) first.</p> <p>3. Run the text processing notebook (code/text_processing.ipynb).</p> <p>4. Run the topic modelling notebook (code/topic_modelling.ipynb).</p> <p>5. Run the topic comparisons notebook (code/topic_comparisons.ipynb).</p> <p>6. Finally, run the analysis notebook (code/analysis/paper_results.ipynb) to get the results as shown on the paper.</p>
Data from: Multispecies site occupancy modeling and study design for spatially replicated environmental DNA metabarcoding
<p>Although environmental DNA (eDNA) metabarcoding has become widely applied to gauge ecosystems in a noninvasive and cost-efficient manner, false negatives can occur due to various factors in its inherent multistage workflow. It is therefore essential to deal with this kind of species detection errors in eDNA metabarcoding to achieve accurate assessment of species distribution and diversity. To address this issue, we proposed a variant of the multispecies site occupancy model for eDNA metabarcoding studies and applied it to an eDNA metabarcoding dataset of freshwater fish communities collected in the Kasumigaura watershed in Japan.</p> <ul> </ul>
Replication Package for the Paper: Understanding Software Architecture Erosion: A Systematic Mapping Study
<p>This is the replication package for the paper: "Understanding Software Architecture Erosion: A Systematic Mapping Study". It contains five files as described below:</p> <p><strong>1. List of the Selected Studies_SMS.xlsx</strong><br> includes the detailed information of the 73 selected studies.</p> <p><strong>2. Data Extraction.xlsx</strong><br> includes the extracted data based on the data items (i.e., D1-D16).</p> <p><strong>3. Sample of Study Selection.xlsx</strong><br> includes the 100 randomly selected papers from the search results and the results of each round selection.</p> <p><strong>4. Pilot Data Extraction.xlsx</strong><br> includes the pilot data extraction results from five papers.</p> <p><strong>5. Sample of Data Extraction.xlsx</strong><br> includes the data extraction results from another randomly selected five papers.</p>
Replication Package for the Paper: "Code Reviewer Recommendation for Architecture Violations: An Exploratory Study"
<p>This is the replication package for the paper: "Code Reviewer Recommendation for Architecture Violations: An Exploratory Study".</p> <p><strong>1) scripts.zip </strong>includes the Python scripts used to run the experiments in this work. Experimental details (e.g., parameters) are described in the Python files. Choose the relevant experimental settings and run "Experiment.py" to start the experiments.</p> <p><strong>2) dataset.xlsx </strong>is the dataset used in the experiments on code reviewer recommendation, which includes the code review comments (from the four OSS projects) related to architecture violations and the file paths of code changes.</p>
Replication package for the Helm charts empirical study
<p><strong>Helm Charts for Kubernetes Applications: Evolution, Outdatedness and Security Risks</strong></p> <p>This repository represents a replication package for our MSR study on Helm charts.</p> <p>This replication package requires Python 3.5+ to be installed, and all the dependencies listed in ``requirements.txt``.</p> <p>They can be automatically installed using ``pip install -r requirements.txt``. <br> These experiments were executed on a Linux Ubuntu OS.</p> <p>This replication package contains three folders:<br> - notebooks: contains notebooks where we analyze data. <br> - figures: contains figures saved from the notebooks<br> - datasets: contains all datasets required</p> <p>To obtain the analysis used in the paper, one should execute ``jupyter notebook`` at the root of this replication package, and open the notebook contained in ``notebooks``.</p> <p>The data is under the Creative Commons Attribution Share-Alike 4.0 license. The source code is under the GNU General Public License.</p>
Replication package for "Method Chaining Redux: An Empirical Study of Method Chaining in Java, Kotlin, and Python"
<p>This dataset includes scripts and data files used to generate all analysis and results from the paper. A <strong>README.md</strong> file is included for details on using the scripts.</p> <p>The dataset is quite large. It is broken down into three archives. All scripts are in <strong>replication-pkg.zip</strong> and the other 2 files only contain data. So if you want to just inspect the analysis, you only need that single zip.</p> <p>If you grab the <strong>data-cached.zip</strong> file and extract it, it will need around 3GB of space. This is the processed dataset stored in Parquet files. Use this if you want to just recreate the tables/figures from the paper.</p> <p>If you want to make changes to the analyses, you will need the raw data in <strong>data-raw.zip</strong>. This will need around 29GB of space once extracted. If you then generate the CSV files from those TXT files (which you will need to do for any custom analysis), you will need an additional 22GB of space.</p>
Replication Package: A Systematic Mapping Study on Security in Configurable Safety-critical Systems Based on Product-Line Concepts
<p><strong>Welcome to the public repository for the additional content of the paper "A Systematic Mapping Study on Security in Configurable Safety-critical Systems Based on Product-Line Concepts", accepted at the ICSOFT 2023.</strong></p> <p>This repository provides additional information to the conducted mapping study, including the following files:</p> <ul> <li>fetched_results_ICSOFT2023.csv: sheet containing all papers fetched from IEEEXplore, Scopus, and the ACM Guide to Computing Literature.</li> <li>excluded_paper.csv: sheet containing all excluded papers related to safety-critical systems but not referring to security.</li> <li>analysis_sheet_ICSOFT2023.csv: sheet containing information regarding the analysis results of 44 included papers based on the extraction criteria.</li> </ul>
Replication package of "How Do Deep Learning Faults Affect AI-Enabled Cyber-Physical Systems in Operation? A Preliminary Study Based on DeepCrime Mutation Operators"
<p>Cyber-Physical Systems (CPSs) combine digital cyber technologies with physical processes. As in any other software system, in the case of CPSs, the use of Artificial Intelligence (AI) techniques in general, and Deep Neural Networks (DNNs) in particular, is contantly increasing. While recent studies have considerably advanced the field of testing AI-enabled systems, it has not yet been investigated how different Deep Learning (DL) bugs affect AI-enabled CPSs in operation. This work-in-progress paper presents a preliminary evaluation on how such bugs can affect CPSs in operation by using a mobile robot as a case study system. For that, we generated DL mutants by using operators proposed by Humbatova et al., which are operators based on real-world DL faults. Our preliminary investigation suggests that such bugs are more difficult to detect when they are deployed in operation rather than when testing their DNN in an off-line setup, which contrast with related studies.</p> <p> </p> <p>This repository provides the replication data employed in our study.</p>
HOPS Study: A Conceptual Replication
ClinicalTrials.gov study NCT04465708. IPD Sharing: NO. Countries: 1. Publications: 6.
Data from: Multispecies site occupancy modeling and study design for spatially replicated environmental DNA metabarcoding
Open the record for dataset details and reuse information.
Evolution repeats itself in replicate long-term studies in the wild
Open the record for dataset details and reuse information.
Does Unit-Tested Code Crash? A Case Study of Eclipse: Replication Package
<p><strong>Does Unit-Tested Code Crash? A Case Study of Eclipse: Replication Package</strong></p> <p>This is a replication package associated with the paper titled “Does Unit-Tested Code Crash? A Case Study of Eclipse”. Below is a description of the package’s contents.</p> <p><strong>Data</strong></p> <p>Data files associated with the paper are provided in the <code>data</code> directory.</p> <p><strong>Text file <code>tested-crashed.txt</code></strong></p> <p>Data specifying whether methods were tested and whether they crashed (according to the criteria adopted in the study). Extracted from <code>matches.xlsx</code>. The data are used as input for Fisher’s test (RQ1).</p> <p><strong>Spreadsheet <code>matches.xlsx</code></strong></p> <p>Test coverage data and calculations associated with failed methods, class coverage, and matching method coverage results are provided in an Excel spreadsheet. Below is the description of the spreadsheet’s contents.</p> <p>Worksheet <em>Test Coverage</em></p> <p>Contains the data regarding the JaCoCo test code coverage analysis.</p> <ul> <li>Class: The name of the class in which a method appears in JVM internal form notation</li> <li>Method: The method’s name</li> <li>Parameters: The method’s arguments in JVM parameter descriptor format; required to handle Java’s {} polymporphism</li> <li>Class Has Unit Test: Whether the corresponding class has associated unit test code</li> <li>Class Unit-Test Line Density: The ratio of lines in class’s test code over those in the class’s implementation code</li> <li>Covered Instructions / Branches / Lines: As reported by JaCoCo</li> <li>Total Instructions / Branches / Lines: As reported by JaCoCo</li> <li>Covered Instructions / Branches / Lines ratio: The ratio between the two preceding values; 1 for methods without any branches</li> <li>Top-1 / Top-6 / Top-10 : In how many stack traces the method appears within; the top-10 / top-6 / the very first stack frame(s)</li> <li>Tested: TRUE if the method is considered tested by having a test code coverage above the median (0.966) and an associated test class</li> <li>Crashed: TRUE if the method has crashed as evidenced by its appearance on the topmost stack frame</li> <li>Stack trace file names: in which the method appeared</li> </ul> <p>Worksheet <em>Test Existence</em></p> <p>Contains the data of the analysis regarding the existence of test code.</p> <ul> <li>Class: Class containing implementation code</li> <li>TestClassNames: Classes that contain tests for the above</li> <li>Number of relevant tests</li> <li>Lines in class test code</li> <li>Lines of class</li> <li>Class Unit-Test Line Density: The ratio between the two above</li> </ul> <p>Worksheet <em>Metrics</em></p> <p>Contains the derivation of metrics reported in the paper. In the cases of tables these are formatted in LaTeX for direct incorporation into the text.</p> <p><strong>Spreadsheet <code>jacoco.xlsx</code></strong></p> <p>Complete test coverage data obtained from JaCoCo are provided in an Excel spreadsheet. Below is the description of the spreadsheet’s contents.</p> <p>Worksheet <em>Data</em></p> <p>Contains the following method code coverage fields as reported by JaCoCo, as well as the calculated percentages.</p> <ul> <li>Class</li> <li>Method</li> <li>Parameters</li> <li>Covered Instructions</li> <li>Total Instructions</li> <li>% Covered Instructions</li> <li>Covered Branches</li> <li>Total Branches</li> <li>% Covered Branches</li> <li>Covered Lines</li> <li>Total Lines</li> <li>% Covered Lines</li> </ul> <p>Worksheet <em>Metrics</em></p> <p>Contains the derivation of numbers reported in the preliminary quantitative analysis and Figure 2.</p> <p>Compressed tar archive <code>eclipse-src.tar.gz</code></p> <p>Contains the Eclipse source code used for running the Eclipse tests with JaCoCo code coverage analysis. It was obtained from the Eclipse source code repositories as follows.</p> <ul> <li>Clone the Eclipse aggreagator repository into a directory named z by running: <code>git clone -b master --recursive git://git.eclipse.org/gitroot/platform/eclipse.platform.releng.aggregator.git z</code></li> <li>In the <code>z</code> directory, checking out the used release by running <code>cd z && git submodule foreach git checkout M20160212-1500</code></li> <li>Checking out the release for the main repository by running: <code>git checkout M20160212-1500</code></li> <li>Applying the patch <code>eclipse-src.diff</code></li> </ul> <p><strong>Patch file <code>eclipse-src.diff</code></strong></p> <p>See above.</p> <p><strong>Zip file <code>incidents.zip</code></strong></p> <p>Contains the 126,026 incidents (crash report stack traces and meta-data) associated with <em>EclipseProduct</em> <code>org.eclipse.epp.package.java.product</code> and <em>BuildID</em> <code>4.5.2.M20160212-1500</code>. This is a subset from the two million incidents available as the <a href="http://software-data.org/datasets/aeri-stacktraces/downloads/incidents_full.tar.bz2">AERI stack traces data set</a>.</p> <p>The subset of incidents was extracted from the full AERI data set with the following command.</p> <pre><code class="language-bash">for f in *; do grep -q org.eclipse.epp.package.java.product $f && grep -q 4.5.2.M20160212-1500 $f && mv $f selected-files/ done</code></pre> <p><br> <strong>Compressed file <code>jacoco.xml.gz</code></strong></p> <p>Contains the results of the JaCoCo code coverage analysis over the Eclipse testing.</p> <p><strong>Code</strong></p> <p>The following scripts are provided in the <code>src</code> directory</p> <ul> <li><code>extract.py</code>: script for extracting crash (incidents) and coverage (JaCoCo) data</li> <li><code>unit-tested-classes.py</code>: script for finding the classes with associated unit test code</li> <li><code>merge.py</code>: script for matching crash (incidents) with coverage (JaCoCo) data</li> <li><code>fisher.r</code>: R script for running Fisher’s test</li> </ul>
Study replicability dataset 3
<p>This dataset contains all data collected to conduct the studies in Chapter 4 "Code comments for defect prediction" of the Ph.D. thesis "Augmented fine-grained defect prediction for code review".</p> <p> </p> <p>Code comments are a key software component containing information about the underlying implementation. Several studies have shown that code comments enhance the readability of the code. Nevertheless, not all the comments have the same goal and target audience. In this paper, we investigate how 14 diverse Java open and closed source software projects use code comments, with the aim of understanding their purpose. Through our analysis, we produce a taxonomy of source code comments; subsequently, we investigate how often each category occurs by manually classifying more than 40,000 lines of code comments from the aforementioned projects. In addition, we investigate how to automatically classify code comments at line level into our taxonomy using machine learning; initial results are promising and suggest that an accurate classification is within reach, even when training the machine learner on projects different than the target one.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.