Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
677
datasets available to search
ShareScore release 0.9.0
Dataset results
677 results for “Replication package”
Replication package
Open the record for dataset details and reuse information.
Replication package: Code Comprehension Confounders: A Study of Intelligence and Personality
<p>Replication package for:</p> <p><em>S. Wagner and M. Wyrich, "Code Comprehension Confounders: A Study of Intelligence and Personality," in IEEE Transactions on Software Engineering, vol. 48, no. 12, pp. 4789-4801, 1 Dec. 2022, doi: 10.1109/TSE.2021.3127131.</em></p> <p>- The `data` folder contains dataset and R analysis script. We recommend calling `setwd()` before running the script contents, so that the dataset can be properly loaded.<br> - the `materials` folder contains the experimental code snippets and the translated task sheets to evaluate code comprehension performance.</p> <p>Please note that the raw data does not contain the complete data set as we only make the data of those participants public that explicitly agreed to it (which applies to 130 of the 135 participants).</p>
The replication package for the paper "Unsupervised Learning of General-Purpose Embeddings for Code Changes"
<p>The replication package for the paper "Unsupervised Learning of General-Purpose Embeddings for Code Changes".</p> <p>To get the detailed instructions, please, consult the README file in the archive.</p>
Replication package for: The Gift of Moving: Intergenerational Consequences of a Mobility Shock
<p>The package contains replication material for Nakamura, E., J. Sigurdsson, J. Steinsson (2021): "The Gift of Moving: Intergenerational Consequences of a Mobility Shock," Review of Economic Studies.</p>
How Tertiary Studies perform Quality Assessment of Secondary Studies in Software Engineering - Replication Package
<p>Replication Package for the paper:</p> <p>D. Costal, C. Farré, X. Franch, C. Quer. 2021. How Tertiary Studies perform Quality Assessment of Secondary Studies in Software Engineering. CIbSE 2021.</p> <p>Please refer to the above paper if you want to cite/use this data.</p>
Replication Package for the Paper: Characteristics and Challenges of Low-Code Development: The Practitioners' Perspective
<p>This is the replication package for the paper: "Characteristics and Challenges of Low-Code Development: The Practitioners’ Perspective". It contains the dataset collected from online developer communities and the data labelling and encoding file of this study. In the meanwhile, we provide a brief description of the files.</p> <p><strong>1. Dataset folder</strong></p> <p>includes the posts collected from Stack Overflow and Reddit.</p> <p><strong>2. Data Labelling & Encoding.mx20</strong></p> <p>is the results of data labelling and encoding that were analyzed by the MAXQDA tool. We extracted data from the posts, labelled them according to the RQs, and encoded the extracted data using Constant Comparison. The file can be opened by MAXQDA 20 or higher versions, which are available at <a href="https://www.maxqda.com/">https://www.maxqda.com/</a> for download. You may also use the free 14-day trial version of MAXQDA 2020, which is available at <a href="https://www.maxqda.com/trial">https://www.maxqda.com/trial</a> for download.</p>
Inclusion and Exclusion Criteria in Software Engineering Tertiary Studies: A Systematic Mapping and Emerging Framework - Replication Package
<p>Replication Package for the paper:</p> <p>D. Costal, C. Farré, X. Franch, C. Quer. 2021. Inclusion and Exclusion Criteria in Software Engineering Tertiary<br> Studies: A Systematic Mapping and Emerging Framework. ESEM '21, <a href="https://doi.org/10.1145/3475716.3484190">https://doi.org/10.1145/3475716.3484190</a></p> <p>Please refer to the above paper if you want to cite/use this data.</p>
Replication Package for "Fertility and Modernity"
<p>This replication package allows replication of all the empirical results in the paper "Fertility and Modernity" by Spolaore and Wacziarg. Please see the readme.pdf file for details.</p>
Replication package for: Forbidden Fruits: The Political Economy of Science, Religion, and Growth
<p>The zipped folder contains all the files need to replicate the results reported in the paper “Forbidden Fruits: The Political Economy of Science, Religion, and Growth” by Roland Bénabou, Davide Ticchi, and Andrea Vindigni, published in the Review of Economic Studies.</p>
Replication package for: "Who Chooses Commitment? Evidence and Welfare Implications"
<p>This package contains all of the code necessary to reproduce the figures and tables in Carrera, Royer, Stehr, Sydnor, Taubinsky (forthcoming) "Who Chooses Commitment? Evidence and Welfare Implications." Review of Economic Studies.</p>
The State of Serverless Applications: Collection,Characterization, and Community Consensus - Replication Package
<p>The replication package for our article <em>The State of Serverless Applications: Collection,Characterization, and Community Consensus</em> provides everything required to reproduce all results for the following three studies:</p> <ul> <li>Serverless Application Collection</li> <li>Serverless Application Characterization</li> <li>Comparison Study</li> </ul> <p><strong>Serverless Application Collection</strong></p> <p>We collect descriptions of serverless applications from open-source projects, academic literature, industrial literature, and scientific computing.</p> <p><em>Open-source Applications</em></p> <p>As a starting point, we used an existing data set on open-source serverless projects from <a href="https://gupea.ub.gu.se/bitstream/2077/62544/1/gupea_2077_62544_1.pdf">this study</a>. We removed small and inactive projects based on the number of files, commits, contributors, and watchers. Next, we manually filtered the resulting data set to include only projects that implement serverless applications. We provide <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Collection/Open%20source%20filtering.xlsx">a table</a> containing all projects that remained after the filtering alongside the notes from the manual filtering.</p> <p><em>Academic Literature Applications</em></p> <p>We based our search on an <a href="https://doi.org/10.5281/zenodo.1175423">existing community-curated dataset</a> on literature for serverless computing consisting of over 180 peer-reviewed articles. First, we filtered the articles based on title and abstract. In a second iteration, we filtered out any articles that implement only a single function for evaluation purposes or do not include sufficient detail to enable a review. As the authors were familiar with some additional publications describing serverless applications, we contributed them to the community-curated dataset and included them in this study. We provide <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Collection/Academic%20literature%20filtering.xlsx">a table</a> with our notes from the manual filtering.</p> <p><em>Scientific Computing Applications</em></p> <p>Most of these scientific computing serverless applications are still at an early stage and therefore there is little public data available. One of the authors is employed at the German Aerospace Center (DLR) at the time of writing, which allowed us to collect information about several projects at DLR that are either currently moving to serverless solutions or are planning to do so. Additionally, an application from the German Electron Synchrotron (DESY) could be included. For each of these scientific computing applications, we provide a document containing a description of the project and the names of our contacts that provided information for the characterization of these applications.</p> <ul> <li>SC1 Copernicus Sentinel-1 for near-real-time water monitoring</li> <li>SC2 Reprocessing Sentinel 5 Precursor data with ProEO</li> <li>SC3 High-Performance Data Analytics for Earth Observation</li> <li>SC4 Tandem-L exploitation platform</li> <li>SC5 Global Urban Footprint</li> <li>SC6 DESY - High Throughput Data Taking</li> </ul> <p><em>Collection of serverless applications</em></p> <p>Based on the previously described methodology, we collected a diverse dataset of 89 serverless applications from open-source projects, academic literature, industrial literature, and scientific computing. This dataset is can be found <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Characterization/Dataset.xlsx">i</a>n Dataset.xlsx.</p> <p><strong>Serverless Application Characterization</strong></p> <p>As previously described, we collected 89 serverless applications from four different sources. Subsequently, two randomly assigned reviewers out of seven available reviewers characterized each application along 22 characteristics in a structured collaborative review sheet. The characteristics and potential values were defined a priori by the authors and iteratively refined, extended, and generalized during the review process. The initial moderate inter-rater agreement was followed by a discussion and consolidation phase, where all differences between the two reviewers were discussed and resolved. The six scientific applications were not publicly available and therefore characterized by a single domain expert, who is either involved in the development of the applications or in direct contact with the development team.</p> <p><em>Initial Ratings & Interrater Agreement Calculation</em></p> <p>The initial reviews are available as <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Characterization/Initial%20Characterizations.csv">a table</a>, where every application is characterized along with the 22 characteristics. A single value indicates that both reviewers assigned the same value, whereas a value of the form <code>[Reviewer 2] A | [Reviewer 4] B</code> indicates that for this characteristic, reviewer two assigned the value A, whereas reviewer assigned the value B.</p> <p>Our script for the calculation of the Fleiß-Kappa score based on this data is also <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Characterization/CalculateKappa.py">publically available</a>. It requires the python package <code>pandas</code> and <code>statsmodels</code>. It does not require any input and assumes that the file <code>Initial Characterizations.csv</code> is located in the same folder. It can be executed as follows:</p> <pre><code>python3 CalculateKappa.py </code></pre> <p><em>Results Including Unknown Data</em></p> <p>In the following discussion and consolidation phase, the reviewers compared their notes and tried to reach a consensus for the characteristics with conflicting assignments. In a few cases, the two reviewers had different interpretations of a characteristic. These conflicts were discussed among all authors to ensure that characteristic interpretations were consistent. However, for most conflicts, the consolidation was a quick process as the most frequent type of conflict was that one reviewer found additional documentation that the other reviewer did not find.</p> <p>For six characteristics, many applications were assigned the ''Unknown'' value, i.e., the reviewers were not able to determine the value of this characteristic. Therefore, we excluded these characteristics from this study. For the remaining characteristics, the percentage of ''Unknowns'' ranges from 0–19% with two outliers at 25% and 30%. These ''Unknowns'' were excluded from the percentage values presented in the article. As part of our replication package, we provide the raw results for each characteristic including the ''Unknown'' percentages in the form of bar charts.</p> <p>The script for the generation of these bar charts is also <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Characterization/GenerateResultsIncludingUnknown.py">part of this replication package</a>). It uses the python packages <code>pandas</code>, <code>numpy</code>, and <code>matplotlib</code>. It does not require any input and assumes that the file <code>Dataset.csv</code> is located in the same folder. It can be executed as follows:</p> <pre><code>python3 GenerateResultsIncludingUnknown.py </code></pre> <p><em>Final Dataset & Figure Generation</em></p> <p>In the following discussion and consolidation phase, the reviewers compared their notes and tried to reach a consensus for the characteristics with conflicting assignments. In a few cases, the two reviewers had different interpretations of a characteristic. These conflicts were discussed among all authors to ensure that characteristic interpretations were consistent. However, for most conflicts, the consolidation was a quick process as the most frequent type of conflict was that one reviewer found additional documentation that the other reviewer did not find. Following this process, we were able to resolve all conflicts, resulting in a collection of 89 applications described by 18 characteristics. This dataset is available here: <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Characterization/Dataset.xlsx">link</a></p> <p>The script to generate all figures shown in the chapter "Serverless Application Characterization can be found <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Serverless%20Application%20Characterization/GenerateFigures.py">here</a>. It does not require any input but assumes that the file <code>Dataset.csv</code> is located in the same folder. It uses the python packages <code>pandas</code>, <code>numpy</code>, and <code>matplotlib</code>. It can be executed as follows:</p> <pre><code>python3 GenerateFigures.py </code></pre> <p><em>Comparison Study</em></p> <p>To identify existing surveys and datasets that also investigate one of our characteristics, we conducted a literature search using Google as our search engine, as we were mostly looking for grey literature. We used the following search term:</p> <pre><code>("serverless" OR "faas") AND ("dataset" OR "survey" OR "report") after: 2018-01-01 </code></pre> <p>This search term looks for any combination of either serverless or faas alongside any of the terms dataset, survey, or report. We further limited the search to any articles after 2017, as serverless is a fast-moving field and therefore any older studies are likely outdated already. This search term resulted in a total of 173 search results. In order to validate if using only a single search engine is sufficient, and if the search term is broad enough, we checked if the seven studies the authors were already familiar with are contained in the search results. As all seven studies were contained in the search results, we concluded that the literature search was broad enough. In a first iteration, we filtered out all results that do not either report original data or report on data from another study. Next, we removed all reports on secondary data, where the original study was already contained in the search results. This process resulted in a total of 16 identified studies. Finally, we determined for each identified study if they investigate one of our characteristics. This resulted in a total of ten related studies. The results from the literature search and the notes from the filtering are part of this replication package as <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Comparison%20Study/ComparisonSearch.xlsx">a table</a>.</p> <p>As these studies use different answer options than our study, we mapped their answer options to ours. In many cases, this was straightforward, such as mapping HTTP to HTTP Request. If the answer option granularities between the studies differed, we aggregated answer options from the study with lower granularity to match the higher granularity study. In case the lower granularity study allowed multiple answers, we selected only the highest value instead of aggregating them, to avoid counting a single study participant multiple times. As this mapping process is somewhat subjective, we provide a detailed account of the mapping for each characteristic and related study <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Comparison%20Study/Comparison%20Mappings.xlsx">as a multi-sheet excel table</a>, where each sheet shows our mapping alongside our notes which answer options were mapped.</p> <p>For many studies, not all information required for traditional meta-analysis techniques, such as cohort size, is available, preventing the application of these meta-analysis techniques. Therefore, we came up with an agreement metric that equally weights the agreement of the reported ranking and the agreement of the reported percentage values. It combines the relative difference between the reported percentages of both studies and the order of the reported popularities of the answer options. We categorize scores in the range [0.8, 1] as very high agreement, [0.6,0.6[ as high agreement, [0.4, 0.6[ as medium agreement, [0.2, 0.4[ as low agreement, and [0, 0.2[ as very low agreement. We acknowledge that these categories are somewhat arbitrary, however, based on a manual inspection of the results, they do seem to capture the level of agreement between the individual studies quite well. Our replication package includes <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Comparison%20Study/Comparison%20Mappings.xlsx">the mapped data</a> alongside the resulting scores to enable a manual inspection of the degree of agreement. The script that implements the calculation of our score is also <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Comparison%20Study/corroboration_analysis.py">publically available</a>. It uses the python packages <code>pandas</code>, <code>numpy</code>, and <code>scipy</code>. The script does not require any input and assumes that the file <code>Comparison Mappings.xlsx</code> is located in the same folder. The script can be executed as follows:</p> <pre><code>python3 corroboration_analysis.py </code></pre> <p>Further, the script to generate all figures shown in the chapter <em>Comparison Analysis</em> is the final piece of our replication package: <a href="https://github.com/ServerlessApplications/ReplicationPackage/blob/main/Comparison%20Study/barcharts.py">link</a>. It uses the python packages <code>numpy</code> and <code>matplotlib</code>. It does not require any input and can be executed as follows:</p> <pre><code>python3 barcharts.py </code></pre> <p>If you have any questions about our study or require any additional information/data please contact the first author.</p>
Replication package for: Geography and Agricultural Productivity: Cross-Country Evidence from Micro Plot-Level Data
<p>The package contains all the data and code necessary to reproduce the results in Adamopoulos and Restuccia (forthcoming). "Geography and Agricultural Productivity: Cross-Country Evidence from Micro Plot-Level Data." <em>Review of Economic Studies</em>. Detailed instructions on how to use the files are provided in a README file.</p>
Replication package for: "Highways, Market Access, and Spatial Sorting".
<ul> <li>The replication package contains data that we are making publicly available.</li> <li>We use two main datasets: census data (Swiss Federal Statistical Office) and taxpayer data (Swiss Federal Tax Administration).</li> <li>Census data are aggregated at the municipality level.</li> <li>Taxpayer data are proprietary. We aggregated these data into four income bins for each municipality. There are fewer than 5 observations in one of these income bins for a bit less than a third of tax data units (municipality-year) and the Swiss Federal Tax Administration asked us to code these bins as “missing” in the dataset we are making publicly available. Hence, the full dataset we are using and the dataset we are making publicly available differ slightly.</li> <li>The ReadMe file includes an Appendix with all Tables and Figures of the paper produced with the restricted, publicly available data.</li> </ul>
Replication package for the paper "Do Comments follow Commenting Conventions? A case study in Java and Python"
<pre><code class="language-markdown"># RP-comment-convention-adherence-Java-Python Replication Package for the paper "Do Comments follow Commenting Conventions? A case study in Java and Python". It uses the dataset provided by Rani et.al.'s work [How to identify class comment types? A multi-language approach for class comment classification](https://github.com/poojaruhal/RP-class-comment-classification). ## Structure ``` RQ1/ RQ1_Java_Rules.xlsx RQ1_Python_Rules.xlsx RQ2/ RQ1_Java_Comments_Validated.xlsx RQ1_Python_Comments_Validated.xlsx Raw-projects/ Java_projects/ eclipse.zip guava.zip guice.zip hadoop.zip spark.zip vaadin.zip Python_projects/ django.zip ipython.zip Mailpile.zip pandas.zip pipenv.zip pytorch.zip requests.zip Style-guides ``` ## Contents of the Replication Package --- - **RQ1/** - contains the data used to answer RQ1 - `RQ1_Java_Rules.xlsx` - contains comment-related rules extracted from various Java style guidelines. Various tabs in the sheet represent the rules extracted from standard or project-specific guidelines. Oracle and Google are the standard guidelines, and the remaining are specific to the projects. - `RQ1_Python_Rules.xlsx` - contains comment-related rules extracted from various Python style guidelines. Various tabs in the sheet represent the rules extracted from standard or project-specific guidelines. PEP, Numpy, and Google are the standard guidelines and the remaining are specific to the projects. - **RQ2/** - contains the data used to answer RQ2 - `RQ2_Java_Comments_Validated.xlsx` - contains Java comment dataset used from the previous work and validated against the rules from their corresponding guidelines. Various tabs in the sheet represent various Java projects used in the work. The rows in each tab show the sample class comments used to validate against the rules. The rules are shown in the columns. - `RQ2_Python_Comments_Validated.xlsx` - contains Python comment dataset used from the previous work and validated against the rules from their corresponding guidelines. Various tabs in the sheet represent various Java projects used in the work. The rows in each tab show the sample class comments used to validate against the rules. The rules are shown in the columns. - **Raw-projects/** contains the raw projects of each language that are used to analyze class comments. - **Java_projects/** - `eclipse.zip` - Eclipse project downloaded from the GitHub. More detail about the project is on https://github.com/eclipse - `guava.zip` - Guava project downloaded from the GitHub. More detail about the project is on https://github.com/google/guava - `guice.zip` - Guice project downloaded from the GitHub. More detail about the project is on https://github.com/google/guice - `hadoop.zip` - Apache Hadoop project downloaded from the GitHub. More detail about the project is on https://github.com/apache/hadoop - `spark.zip` - Apache Hadoop project downloaded from the GitHub. More detail about the project is on https://github.com/apache/spark - `vaadin.zip` - Vaadin project downloaded from the GitHub. More detail about the project is on https://github.com/vaadin/framework - **Python_projects/** - `django.zip` - Django project downloaded from the GitHub. More detail about the project is on https://github.com/django. - `ipython.zip` - IPython project downloaded from the GitHub. More detail about the project is on https://github.com/ipython/ipython - `Mailpile.zip` - Mailpile project downloaded from the GitHub. More detail about the project is on https://github.com/mailpile/Mailpile - `pandas.zip` - pandas project downloaded from the GitHub. More detail about the project is on https://github.com/pandas-dev/pandas - `pipenv.zip` - Pipenv project downloaded from the GitHub. More detail about the project is on https://github.com/pypa/pipenv - `pytorch.zip` - PyTorch project downloaded from the GitHub. More detail about the project is on https://github.com/pytorch/pytorch - `requests.zip` - Requests project downloaded from the GitHub. More detail about the project is on https://github.com/psf/requests/ - **Style-guides/**- contains the style guidelines used for the selected projects. ---</code></pre> <p> </p>
Replication package for Seawalls and Stilts: A Quantitative Macro Study of Climate Adaptation
<p>This replication package contains the data sets and code necessary to reproduce the results in Fried, Stephie. "Seawalls and Stilts: A Quantitative Macro Study of Climate Adaptation." accepted for publication at the Review of Economic Studies. </p> <p> </p>
Replication package for "Exporting and Offshoring with Monopsonistic Competition"
<p>The code in this replication package constructs Figure A.1 and the data reported in Table A.1 of the paper "Exporting and Offshoring with Monopsonistic Competition" (by Egger, H., Kreickemeier, U., Moser, C. and Wrona, J.)</p>
Replication package for: Entrepreneurial Human Capital and Firm Dynamics
<p>This replication package contains publicly available data and code to replicate the figures and tables in:</p> <p>Queiró, Francisco. 2021. Entrepreneurial Human Capital and Firm Dynamics.</p>
Replication package for "Real Effects of Financial Distress: The Role of Heterogeneity"
<p>This package contains synthetic data and code that generates the figures and tables in the paper and the online appendix. Data files in .csv and stata format have been provided. A readme has been included that provides detailed guidelines how to run the codes.</p>
Replication Package: Where to Go Now? Finding Alternatives for Declining Packages in the npm Ecosystem
<p>Software ecosystems are the backbone of modern software developments, which make it grow exponentially. Developers add new packages every day to solve new problems or provide alternative solutions, causing obsolete packages to decline in their importance to the community. Packages in decline are reused less over time and may become less frequently maintained. Thus, developers usually migrate their dependencies to better alternatives. Replacing packages in decline with better alternatives requires time and effort by developers to allocate packages that need to be replaced, find the alternatives, asset migration benefits, and finally, perform the migration. This paper proposes an approach that automatically identifies packages that need to be replaced and find their alternatives supported with real-word example of open source projects performing the suggested migration. At its core, our approach relies on the dependency migration patterns performed in the ecosystem to suggest migrations to other developers. We evaluated our approach on the npm ecosystem and found that 96% of the suggested alternatives are accurate. Furthermore, by surveying expert JavaScript developers, 67% of them indicate that they will use our suggested alternative packages in their future projects.</p>
Replication Package for: "Pre-Colonial Warfare and Long-Run Development in India"
<p>Data and code to replicate all the results in the article Dincecco, Mark; Fenske, James; Menon, Anil; and Mukherjee, Shivaji "Pre-Colonial Warfare and Long-Run Development in India," Economic Journal.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.