Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
184
datasets available to search
ShareScore release 0.9.0
Dataset results
184 results for “Replication Study”
Source Data for Published Study "Are changes in nociceptive withdrawal reflex magnitude a viable central sensitization proxy? Implications of a replication attempt"
<p>Upload version NWR_v01_20230409</p> <p>Authors: Alexandros Guekos, Alince Catrine Grata, Michèle Hubli, Martin Schubert, and Petra Schweinhardt</p> <p>The present data was collected from August to October 2019 as part of a replication attempt of a previously published study (Ellrich, J., and R-D. Treede. "Convergence of nociceptive and non-nociceptive inputs onto spinal reflex pathways to the tibialis anterior muscle in humans." Acta physiologica scandinavica 163.4 (1998): 391-401, https://doi.org/10.1046/j.1365-201X.1998.t01-1-00392.x). </p> <p>The results of the replication study have been published under open access (Guekos, A., et al. "Are changes in nociceptive withdrawal reflex magnitude a viable central sensitization proxy? Implications of a replication attempt" Clinical Neurophysiology 145 (2023): 139-150, https://doi.org/10.1016/j.clinph.2022.09.011).</p> <p>Details of the paradigm, the experimental setup, and the analysis can be found there.</p> <p>In brief, 16 healthy adults (8 men and 8 women) underwent a single experimental session during which a tonic heat stimulus was applied on one leg to the foot sole and on the other to the calf muscle. Both legs were tested consecutively in pseudorandom order. Concurrently, subjects received transcutaneous electrical stimuli to elicit the nociceptive withdrawal reflex (NWR). The muscle responses were recorded via surface electromyography (sEMG) from the biceps femoris (BF), rectus femoris (RF), and tibialis anterior (TA).</p> <p>The protocol consisisted of eight blocks per leg. During the first two blocks no temperature stimulation was applied. These two blocks served to identify the NWR threshold at the BF. For threshold determination, a single ascending staircase with either single electrical stimulations or triplets (at 2Hz) were used. From the triplets, only the muscle response to the third stimulation was analysed. The higher of the two obtained currents was used as the threshold. The following six blocks used six different temperatures (one per block) of 32, 36, 39, 42, 45 and 46 centigrade. During each block eight transcutaneous electrical stimuli were applied, either to the medial plantar nerve (MP) on the foot sole or to the retromalleolar pathway of the sural nerve (SU). The stimulations increased from -4 mA w.r.t. threshold to 200% threhold. Participants verbally rated perceived pain for every stimulation during these six blocks.</p> <p>Every electrical stimulation consisted of a train of five rectangular stimuli of 1 ms duration delivered at 200 Hz. Muscle responses were recorded from 120 pre- to 380 ms post-stimulation. The recorded sEMG signals were sampled at 48 kHz and downsampled to 6 kHz, rectified, band-pass filtered from 10 Hz to 500 Hz and amplified up to 125 times. Between 120 ms pre- and 380 ms post-stimulation, traces for all applied stimulations were automatically saved into separate txt files.</p> <p>Please consult the README.txt file for details on the structure of the uploaded data and for information w.r.t. potential instances of incompleteness or unusability.</p> <p>The study was funded by the Swiss National Science Foundation as part of a grant to PS (grant number 320030_179191/1).</p>
Penothypic integration: assessing the value of study replication
<p>This repository contains the databases and scripts used to analyze the phenotypic integration of different species, populations, and sexes and assess which structural paths were generally (vs. conditionally) supported (vs. unsupported). These data were used in the following study:</p> <p>Irene Gaona-Gordillo, Benedikt Holtmann, Alexia Mouchet, Alexander Hutfluss, Alfredo Sánchez-Tójar, and Niels J. Dingemanse. <em>Unpublished manuscript. </em>Are animal personality, body condition, physiology, and structural size integrated? A comparison of species, populations, and sexes, and the value of study replication. J Anim Ecol.</p> <p>For any further information, please contact: </p> <p>Irene Gaona-Gordillo, email: gaona-gordillo@bio.lmu.de</p> <p>Niels Dingemanse, email: n.dingemanse@lmu.de</p>
Empirical Study on Test Generation Using GitHub Copilot --- Replication Package
<p>This replication package contains the data and scripts used in the "Empirical Study on Test Generation Using GitHub Copilot" thesis. </p>
Replication Package for "Why Do Deep Learning Projects Differ in Compatible Framework Versions? An Exploratory Study"
<p>This dataset contains scripts and data used to generate relevant results for this paper. Detailed information are described in README.md. </p> <p>code</p> <p>This folder contains all the scripts used for the experiment. The upgrade.py and downgrade.py are used to perform upgrade and downgrade runs. The pairing.py is used to generate the DFVC pairs. The main.py is used to identify root causes of DFVC pairs.</p> <p>result</p> <p>This folder contains all the results of the experiments, including the runtime output (e.g., a_1.0.0.txt), the runtime environment (e.g., condalist_1.0.0.txt), and the project's runtime commands (e.g., pytorch-cifar.xlsx) of all tested 90 PyTorch and 50 TensorFlow projects.</p> <p><br> Distribution of dfvc pairs.xlsx</p> <p>This file includes 6,926 DFVC pairs and their root causes.</p> <p>Tested framework versions.xlsx</p> <p>This file includes the framework versions tested and the Python versions that the framework versions are compatible with.</p> <p>Tested projects.xlsx</p> <p>This file includes the tested 90 PyTorch projects and 50 TensorFlow projects. We provide the following main information: (a) project name, (b) stars, (c) link, (d) the starting version, (e) python version, (f) incompatible upgrade/downgrade version, and (g) compatible versions.</p>
Reproducible Validation and Replication Studies in Nanoscale Physics (repro results plots - Ellis et al., 2016)
<p>This archive contains the Jupyter notebooks needed to reproduce the figures of the paper that are related to the Validation results and replication of Ellis et al. 2016. For further information direct to the README.md file.</p>
Reproducible Validation and Replication Studies in Nanoscale Physics (problem datasets for Rockstuhl et al. 2005 replication)
<p>Problem folders including all the input files necessary to reproduce the computations of the results related to Rockstuhl et al. 2005 on the paper: Reproducible Validation and Replication Studies in Nanoscale Physics</p>
Replication Package for the Paper: "Will Data Influence the Experiment Results?: A Replication Study of Automatic Identification of Decisions"
<p>This is the replication package for the paper: "Will Data Influence the Experiment Results?: A Replication Study of Automatic Identification of Decisions". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package below.</p> <p><strong>1. main_code folder</strong></p> <ul> <li><em>automatic_approach.py </em>contains the main source code of the automatic approach for identifying decisions in our experiment, which is conducted on MacOs and Python 3.7.9. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirement.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirement.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>EASE2020 - 650 decisions.xlsx </em>contains 650 decision sentences from our previous work (EASE2020)</li> <li><em>EASE2020 - 650 non-decisions.xlsx </em>contains 650 non-decision sentences from our previous work (EASE2020)</li> <li><em>Our 844 relabeled decisions.xlsx</em> contains 844 relabeled decisions in this work.</li> <li><em>Our 750 assumptions.xlsx</em> contains 750 assumptions from our previous work (APSEC2019)</li> </ul> <p><strong>3. RQ1 folder</strong></p> <ul> <li><em>experiment_RQ1.py</em> contains the main source code of the experiment for answering RQ1, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul> <p><strong>4. RQ2 folder</strong></p> <ul> <li><em>experiment_RQ2.py</em> contains the main source code of the experiment for answering RQ2, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul> <p><strong>5. RQ3 folder</strong></p> <ul> <li><em>experiment_RQ3.py</em> contains the main source code of the experiment for answering RQ3, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul>
Replication Package for the Paper: "Understanding Code Smell Detection via Code Review: A Study of the OpenStack Community"
<p>This repository contains the data and results from the paper "Understanding Code Smell Detection via Code Review: A Study of the OpenStack Community" submitted to ICPC 2021.</p> <p> </p> <p><strong>1. "data.zip" contains the following three folders:</strong></p> <p> </p> <p><strong>1) data folder</strong></p> <p>The data folder contains the retrieved 1,190 reviews that discuss code smells. Each review includes four parts: Code Change URL, Code Smell, Code Smell Discussion, and Source Code URL.</p> <p> </p> <p><strong>2) scripts folder</strong></p> <p>The scripts folder contains the Python scripts that were used to search for code smell terms and the list of code smell terms.</p> <ul> <li> <p><em>keyword.txt</em> contains the keywords associated with code smells, such as "smell, duplication, and dead".</p> </li> <li> <p><em>get_changes.py</em> is used for getting code changes from OpenStack.</p> </li> <li> <p><em>get_comments.py</em> is used for getting review comments for each code change.</p> </li> <li> <p><em>keywords_search.py</em> is used for searching review comments that contain at least one keyword.</p> </li> <li> <p><em>random_select.py</em> is used for randomly selecting review comments that do not contain any keyword.</p> </li> <li> <p><em>keywords_improve.py</em> is used for improving the keyword-based mining approach.</p> </li> <li> <p><em>tools.py</em> is used for supporting the process of keywords improving.</p> </li> </ul> <p> </p> <p><strong>3) project folder</strong></p> <p>The project folder contains the MAXQDA project files. The files can be opened by MAXQDA 12 or higher versions, which are available at <a href="https://www.maxqda.com/">https://www.maxqda.com/</a> for download. You may also use the free 14-day trial version of MAXQDA 2018, which is available at <a href="https://www.maxqda.com/trial">https://www.maxqda.com/trial</a> for download.</p> <ul> <li> <p><em>Data Labeling & Encoding for RQ2.mx12</em> is the results of data labeling and encoding for RQ2, which were analyzed by the MAXQDA tool.</p> </li> <li> <p><em>Data Labeling & Encoding for RQ3.mx12</em> is the results of data labeling and encoding for RQ3, which were analyzed by the MAXQDA tool.</p> </li> </ul> <p> </p> <p><strong>2. Keywords associated with code smells.pdf</strong></p> <p>This file contains the final set of keywords associated with code smells that we identified by following the systematic approach proposed by Bosu and his colleagues in their paper: Identifying the Characteristics of Vulnerable Code Changes: An Empirical Study, FSE 2014.</p>
Replication package for How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists
<p>This dataset was used in the paper: "How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists", Journal of Empirical Software Engineering, to appear.</p>
Replication package for the paper :The Relationship Between Different Python Argument-Passing Mechanisms and Fixes: An Empirical Study
<p><strong>Abstract:</strong></p> <p>Modern programming languages, such as Python, have introduced a variety of constructs and syntactical elements to make software development more efficient and concise. Examples include lambda functions, comprehension collections, or mechanisms to facilitate the passing of arguments to a function. While many of such constructs may, in principle, be beneficial for developers, recent studies have shown that certain programming constructs may affect program understanding and even induce more fixes than other changes. <br>This paper studies the effect of different Python argument-passing mechanisms to investigate their relationship with code proneness to be fixed. Specifically, we study the fix-proneness for what concerns function definitions and invocations. This is done by analyzing the evolutionary history of 200 Python projects, for a total of about 3M functions and 12M call sites. While there are varying effects for what concerns parameter declaration mechanisms, we found evidence that keyword-based argument passing is less defect-prone than positional argument passing, and this is not affected by size-related confounding factors.</p>
Replication Package for "Benefits and pitfalls of token-level SZZ: An empirical study on OSS projects"
<p>Replication Package for "Benefits and pitfalls of token-level SZZ: An empirical study on OSS projects"</p><p>All materials are licensed under the MIT License (see LICENSE file). </p>
Replication Package: Product-Line Engineering for Smart Manufacturing: A Systematic Mapping Study on Security Concepts
<p><strong>Welcome to the public repository for the additional content of the paper "Product-Line Engineering for Smart Manufacturing: A Systematic Mapping Study on Security Concepts", accepted at the ICSOFT 2024.</strong></p> <p>This repository provides additional information to the conducted mapping study, including the following file:</p> <ul> <li>analysis_sheet_ICSOFT2024.csv: sheet containing information regarding the analysis results of 43 included papers based on the extraction criteria.</li> </ul>
Model Generation from Requirements with LLMs: an Exploratory Study - Replication Package
<p>This is a replication package for the paper "<span>Model Generation from Requirements </span><span>with LLMs: an Exploratory Study</span>", by Sallam Abualhaija, Chetan Arora, and Alessio Ferrari.</p> <p><strong>Abstract: </strong>Complementing natural language (NL) requirements with graphical models can improve stakeholders’ communication and provide directions for system design. However, creating models from requirements involves manual effort. The advent of generative large language models (LLMs), ChatGPT being a notable example, offers promising avenues for automated assistance in model generation. This paper investigates the reliability of ChatGPT in generating sequence diagrams from NL requirements. Specifically, we conduct a qualitative study examining the sequence diagrams generated by ChatGPT for 28 requirements documents of various types and from different domains. Our study aims to uncover potential issues that emerge in the models generated by ChatGPT, thereby hindering its applicability in practice. Observations have systematically been captured through evaluation logs, and categorized through thematic analysis. Our results indicate that, although the models generally conform to the standard and exhibit a reasonable level of understandability, their correctness with respect to the specified requirements often presents challenges. This issue is particularly pronounced in the presence of requirements smells, such as ambiguity and inconsistency. The insights derived from this study can influence the practical utilization of LLMs in the RE process, and open the door to novel RE-specific prompting strategies targeting effective model generation.</p> <p>The replication package consists of the following folders:</p> <p><strong>logs:</strong> includes the evaluation logs produced by each evaluator</p> <p><strong>original-documents: </strong>includes the original requirements documents used for the evaluation</p> <p><strong>RQ1 - quantitative analysis:</strong> includes the analysis made on the scores given to each model and model variant. It includes five files:</p> <p>- results.csv: numerical results of the evaluation for each criterion<br>- analysis-results.Rmd: R file used to perform the quantitative analysis (requires R Studio to be executed)<br>- analysis-results.html: html file produced by analysis-results.Rmd<br>- cross-check.csv: file with the cross-checking of the two assessors applied to a subset of the models<br>- symmary_results.xlsx: final output of the quantitative results in terms of Wilcoxon signed rank tests</p> <p><strong>RQ2 - thematic analysis: </strong>includes the codebook produced by the thematic analysis of the issues in generating models with ChatGPT</p>
Policy Testing with MDPFuzz (Replicability Study): RQ2&3
<p>Data and figures used in the replication study (RQ2&3) of the paper <em>Policy Testing with MDPFuzz (Replicability Study)</em>.</p>
Policy Testing with MDPFuzz (Replicability Study): RQ1
<p>Data and figure used in the reproduction study (RQ1) of the paper <em>Policy Testing with MDPFuzz (Replicability Study)</em>.</p>
Revisiting the Building of Past Snapshots – A Replication and Reproduction Study
<p>This dataset contains all the data that could not be included in the original repository: https://github.com/BuildabilityResearcher/BuildabilityStudy</p>
Replication Package for the Paper: Transfer Learning with Time Series Data: A Systematic Mapping Study
<p>This is a replication package for the paper "Transfer Learning with Time Series Data: A Systematic Mapping Study".</p> <p>It provides</p> <ul> <li>a documentation of the conducted electronic literature search,</li> <li>exports of the search results from each literature database,</li> <li>and an excel file on the included literature and extracted data.</li> </ul>
Supplementary Material for "Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies"
<p>This supplementary material for the article"Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies. Dell’Anna, D.; Aydemir, F. B.; and Dalpiaz, F. Empirical Software Engineering. 2022" includes</p> <ul> <li>ECSER-ExploratoryStudy.csv: The annotated meta-data of the papers that have been published in ICSE between 2019 and 2021.</li> <li>ECSER_ROCplots+StatTest.ipynb: A python notebook that compares classifiers adn checks the statistical significance of the comparison results.</li> <li>ECSER_SummaryOfReplicationSteps.pdf: This table presents a summary of ECSER steps for the two original studies and our applications on ECSER.</li> <li>ECSER_RE: The directory that holds the datasets and code for the replication of Hay et al. [1] and additional runs on the new data sets.</li> <li>ECSER_FF: The directory that holds the code and data for the replication of Alshammari et al. [2]</li> <li>README.md presents the structure of the supplementary materials.</li> <li>requirements.txt lists the dependencies needed to run the code</li> </ul> <p>In the ECSER_RE directory, the code for multiple classifiers that are compared are kept in the "Classifiers" directory. The public data sets are shared in the Datasets directory. "ECSER_RE_Compare_Classifiers.ipynb" python notebook includes the code that runs each classifier. The results are presented in "ECSER_RE_results-Promise-vs-all.csv".</p> <p>In the ECSER_FF directory, the data sets are presented directly under the main directory. The notebook "ECSER-FF-Compare_Classifiers.ipynb" compares the classifiers of the original study and the results are kept under the "ECSER_FF_results" directory.</p> <p><strong>How to cite this repository </strong><br>If you use this repository, please cite the reference paper, and the repository, as below:</p> <p>Dell’Anna, Davide, Fatma Başak Aydemir, and Fabiano Dalpiaz. "Evaluating classifiers in SE research: the ECSER pipeline and two replication studies." Empirical Software Engineering 28.1 (2023): 3.</p> <p>Davide Dell'Anna, Fatma Başak Aydemir, & Fabiano Dalpiaz. (2021). Supplementary Material for "Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies" [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6266675</p> <p> </p> <p> </p> <p> </p> <ol> <li>Tobias Hey, Jan Keim, Anne Koziolek, and Walter F. Tichy. 2020. SupplementaryMaterial of "NoRBERT: Transfer Learning for Requirements Classification". https://doi.org/10.5281/zenodo.3874137</li> <li>Abdulrahman Alshammari, Christopher Morris, Michael Hilton, and JonathanBell. 2021. FlakeFlagger: Predicting Flakiness Without Rerunning Tests. In43rdIEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid,Spain, 22-30 May 2021. IEEE, 1572–1584. https://doi.org/10.1109/ICSE43902.2021.00140</li> </ol>
Replication Package for the Paper: "Understanding Code Snippets in Code Reviews: A Preliminary Study of the OpenStack Community"
<p>This is the replication package for the paper: "Understanding Code Snippets in Code Reviews: A Preliminary Study of the OpenStack Community", including dataset and so on (see the description below) : </p> <ul> <li> <p><strong>Data of Code Snippets in Code Review.xlsx</strong> is the dataset of our paper, which contains 10,790 review comments collected from the Nova project and Neutron project of OpenStack community. Among all the review comments, 626 review comments contain code snippets. For the rows of review comments with code snippets, we filled them with blue color as an indicator.</p> </li> <li> <p><strong>Examples for Each Purpose.xlsx</strong> contains six review comment examples for the six detailed purposes mentioned in our paper (see Section 4.2).</p> </li> <li> <p><strong>README.md</strong></p> </li> </ul>
Replication Package for the Paper: "Code Smells Detection via Modern Code Review: A Study of the OpenStack and Qt Communities"
<p>This repository contains the data and results from the paper "Code Smells Detection via Modern Code Review: A Study of the OpenStack and Qt Communities" submitted to the ICPC 2021 special issue of the Empirical Software Engineering Journal, 2021.</p> <p> </p> <p>The replication package contains the following two folders:</p> <p> </p> <p><strong>1) data folder</strong></p> <p>The data folder contains the following four folders, which is organized by research questions (RQs).</p> <ul> <li>RQ1: The RQ1 folder contains the retrieved 1,539 code reviews that discuss code smells. Each review includes four parts: Code Change URL, Code Smell, Code Smell Discussion, and Source Code URL.</li> <li>RQ2: The RQ2 folder contains the coded data for RQ2, called <em>Data Labeling & Encoding for RQ2.mx18</em>. It is the results of data labeling and encoding for RQ2, which was analyzed by the MAXQDA tool.</li> <li>RQ3 and RQ5: <ul> <li><em>Extracted data for RQ3.1.xlsx</em>: this file contains the extracted data (i.e., specific refactoring actions suggested by reviewers) for RQ3.1.</li> <li><em>Data Labeling & Encoding for RQ3 and RQ5.mx18</em>: this file contains the extracted data for RQ3 (excluding the specific refactoring actions in RQ3.1) and RQ5.</li> <li><em>Code change status for RQ5.xlsx</em>: this file contains the information of status of code changes where the developers disagreed with the reviewers and chose to ignore the identified code smells.</li> </ul> </li> <li>RQ4: The RQ4 folder contains the extracted data for RQ4, called <em>Extracted data for RQ4.xlsx</em>.</li> </ul> <p>Note: The mx18 files can be opened by MAXQDA 18 or higher versions, which are available at https://www.maxqda.com/ for download. You may also use the free 14-day trial version of MAXQDA 2018, which is available at https://www.maxqda.com/trial for download.</p> <p> </p> <p><strong>2) scripts folder</strong></p> <p>The scripts folder contains the Python scripts that were used to search for code smell terms and the list of code smell terms.</p> <ul> <li><em>keyword.txt</em> contains the keywords associated with code smells, such as "smell, duplication, and dead".</li> <li><em>get_changes.py</em> is used for getting code changes from OpenStack and Qt.</li> <li><em>get_comments.py</em> is used for getting review comments for each code change.</li> <li><em>keywords_search.py</em> is used for searching review comments that contain at least one keyword.</li> <li><em>random_select.py</em> is used for randomly selecting review comments that do not contain any keyword.</li> <li><em>keywords_improve.py</em> is used for improving the keyword-based mining approach.</li> <li><em>tools.py</em> is used for supporting the process of keywords improving.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.