Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

677

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

677 results for “Replication package”

Learn how ShareScore rates datasets ↗
zenodo32/100

Replication package for Kinship taxation as an impediment to growth: Experimental evidence from Kenyan microenterprises

<p>Munir Squires (2024): &ldquo;Kinship taxation as an impediment to growth: Experimental evidence from Kenyan microenterprises,&rdquo; <em>The Economic Journal</em></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Replication package for: Shallow Meritocracy

<p>The package contains all the code and data necessary to reproduce the figures and tables in Andre (forthcoming), "Shallow Meritocracy", Review of Economic Studies.</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Replication package for: "Structural Change, Elite Capitalism, and the Emergence of Labor Emancipation"

<p>Replication package for "Structural Change, Elite Capitalism, and the Emergence of Labor Emancipation" (Review of Economic Studies) by Quamrul H. Ashraf, Francesco Cinnirella, Oded Galor, Boris Gershman, and Erik Hornung. This package contains data and code for reproducing the tables and figures in the article and its online appendix. Detailed instructions for accessing the raw data are also provided.</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Replication Package for "Trust and State Effectiveness: The Political Economy of Compliance"

<p>There is a README file that provides information on the data used in the paper titled &ldquo;Trust and State Effectiveness: The Political Economy of Compliance&rdquo; by Tim Besley and Sacha Dray, and steps to replicate its results.<br>DATA SOURCES:<br>Raw datasets needed for replication are in the folder "data/raw".<br>1/ IVS_clean.dta : dataset based on the Integrated Values Survey 1981-2021, which can be accessible from the World Values Survey or European Values Survey website.<br>-<br>EVS (2022): EVS Trend File 1981-2017. GESIS Data Archive, Cologne. ZA7503 Data file Version 3.0.0, doi:10.4232/1.14021.<br>-<br>2/ UK_C19_cohorts: dataset from the COVID-19 Survey in Five National Longitudinal Cohort Studies.<br>-<br>Centre for Longitudinal Studies. (2020). COVID-19 survey in five national longitudinal cohort studies: Millennium Cohort Study, Next Steps, 1970 British Cohort Study and 1958 National Child Development Study, 2020&ndash;2021 [data collection].</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Replication Package for Gang Rule: Understanding and countering criminal governance

<p>Blattman, Christopher and Duncan, Gustavo and Lessing, Benjamin and Tobon, Santiago. (Forecoming). "Gang Rule: Understanding and Countering Criminal Governance". The Review of Economic Studies</p> <p>The package contains all the code and data necessary to reproduce the figures and tables in the aforementioned paper.&nbsp;All raw data that can be disclosed was collected by the authors, and are available under a Creative Commons Non-commercial license.&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Replication package for the paper "Modeling Europe's role in the global LNG market 2040: balancing decarbonization goals, energy security, and geopolitical tensions"

<p>This package contains folders and files with code and data used in the study described in the paper. Further information can be found on <a href="https://github.com/sebastianzwickl/lng-trade-europe" target="_blank" rel="noopener">GitHub</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Replication package: Temperature and decision-making

<p>Michelle Escobar Carias, David Johnston, Rachel Knott, Rohan Sweeney (2024). Replication package: Temperature and decision-making</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Replication package for: "Observed Patterns of Free-Floating Car-Sharing Use"

<p>Replication packet for &ldquo;Observed Patterns of Free-Floating Car-Sharing Use.&rdquo; By Natalia Fabra, Catarina Pintassilgo, and Mateus Souza. SERIEs: Journal of the Spanish Economic Association.<br><br>The code in this replication packet generates all the tables and figures from the paper. The car-sharing trips data were obtained under confidentiality agreements thus cannot be publicly shared. To obtain permissions to access the car-sharing data, replicators may contact the authors. All other data are publicly available and included in the packet.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Replication package for: "The Impact of Online Competition on Local Newspapers: Evidence from the Introduction of Craigslist"

<p>Djourelova, Milena, Durante, Ruben, and Gregory J. Martin (2023). "The Impact of Online Competition on Local Newspapers: Evidence from the Introduction of Craigslist."</p><p>The replication package contains software and datasets needed to produce tables and figures in the paper, as well as data cleaning scripts and raw data.</p><p>The replication package is also available as a Git repository at <a href="https://code.stanford.edu/gjmartin/craigslist-replication-code-and-data">https://code.stanford.edu/gjmartin/craigslist-replication-code-and-data</a></p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Replication package for "Does Inequality Affect the Needs of the Poor?"

<p>This package contains the programs and instructions to replicate manuscript &ldquo;Does Inequality Affect the Needs of the Poor?&rdquo; by Eve Colson-Sihra and Cl&eacute;ment S. Bellet forthcoming at JEEA.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Replication package for "The Effect of Asset Encumbrance on Bank Behavior: Evidence from the Introduction of Covered Bonds in Norway"

<p>This package contains <strong>synthetic </strong>data, programs and instructions to replicate manuscript "The Effect of Asset Encumbrance on Bank Behavior:&nbsp;Evidence from the Introduction of Covered Bonds in Norway"&nbsp; by Cao, Juelsrud and Sondershaus, forthcoming at JEEA.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Replication package for the paper: Harnessing Test Call Structures for Improved Fault Localization Effectiveness

<p>This repository contains the replication package for the paper&nbsp;<strong>"Harnessing Test Call Structures for Improved Fault Localization Effectiveness"</strong>.&nbsp;The package includes the necessary scripts,&nbsp;data,&nbsp;and instructions to reproduce the results presented in the study,&nbsp;focusing on the effectiveness of Spectrum-Based Fault Localization&nbsp;(SBFL)&nbsp;using the Barinel algorithm.</p> <h3>Description of Folders and Files:</h3> <ul> <li><strong>algorithms/</strong>: Contains the SBFL algorithms implemented for this study. Note that only the Barinel algorithm was used in the presented results.</li> <li><strong>base/</strong>: Contains the scripts needed for calculating spectra for SBFL.</li> <li><strong>D4J/</strong>: Contains the Defects4J projects. These need to be unpacked using the appropriate scripts provided in this folder.</li> <li><strong>changeset.json</strong>: Contains metadata about changesets for the projects analyzed in this study.</li> <li><strong>main.py</strong>: The main script for calculating ranks and metrics for the selected projects and heuristics.</li> <li><strong>SBFL ranks.csv</strong>: Contains the SBFL ranks calculated from the analysis.</li> </ul> <h2>Requirements</h2> <p>To run the scripts,&nbsp;you will need:</p> <ul> <li><strong>Python 3.9</strong></li> <li><strong>Defects4J</strong>: The&nbsp;<code>D4J</code>&nbsp;folder should contain unpacked Defects4J projects. Please ensure you have followed the instructions in the&nbsp;<code>D4J</code>&nbsp;folder to unpack the projects accordingly.</li> <li><strong>Required Python packages</strong>&nbsp;(install using&nbsp;<code><a href="http://localhost:63343/markdownPreview/1636345610/markdown-preview-index-2076127249.html?_ijt=mqjd55qrfb1m1q15g8bm74pk1u#"></a>pip</code>):&nbsp;<code><a href="http://localhost:63343/markdownPreview/1636345610/markdown-preview-index-2076127249.html?_ijt=mqjd55qrfb1m1q15g8bm74pk1u#"></a>pip install -r requirements.txt</code></li> <li><strong>Prepare the Defects4J Projects</strong>: <ul> <li>Navigate to the&nbsp;<code>D4J</code>&nbsp;directory.</li> <li>Run the provided scripts to unpack the necessary Defects4J projects.</li> <li>Ensure that each project folder is structured properly to be used by the&nbsp;<code>main.py</code>&nbsp;script.</li> </ul> </li> </ul> <h2>Running the Main Script</h2> <p>The main analysis script is&nbsp;<code>main.py</code>,&nbsp;which calculates the SBFL ranks using the Barinel algorithm based on the selected heuristics.</p> <h3>Usage</h3> <p>To run the main script:</p> <div></div> <pre><code><a href="http://localhost:63343/markdownPreview/1636345610/markdown-preview-index-2076127249.html?_ijt=mqjd55qrfb1m1q15g8bm74pk1u#"></a>python main.py </code></pre> <h3>Parameters and Settings:</h3> <ul> <li><strong>Projects and Ranges</strong>: The projects are defined in the&nbsp;<code>projects</code>&nbsp;list, with their respective bug ranges in the&nbsp;<code>ranges</code>&nbsp;list. The script iterates over these projects and bug IDs to calculate SBFL ranks.</li> </ul> <h3>Outputs:</h3> <ul> <li>The script outputs the ranks and coverage metrics directly to the console. Results can be redirected or saved as needed.</li> <li>It uses the&nbsp;<code>Ranks.RankContainer</code> class to calculate and print the minimum suspiciousness ranks for the selected metrics.</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Replication Package for "Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets"

<p>Replication Package for "Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets"<br><br>Includes datasets, and bash scripts.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Replication package for: The Long-Run Labor Market Effects of the Canada-U.S. Free Trade Agreement

<p>Replication package for: The Long-Run Labor Market Effects of the Canada-U.S. Free Trade Agreement, published in the Review of Economic Studies. The replication package includes all publicly available data and all code required to replicate the results in the paper.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Replication Package of the Paper "How do Papers Make into Machine Learning Frameworks: A Preliminary Study on TensorFlow"

<p>This replication package contains datasets and scripts related to the paper: "<em>How do Papers Make into Machine Learning Frameworks: A Preliminary Study on TensorFlow</em>"</p> <ul> <li> <p><code>Contributor_Classification.csv</code>: contains the assignment of each contributor to a specific classification. The file contains the following columns:</p> <ul> <li><em>Date</em> : contains the date of each comment</li> <li><em>Type</em>: describe the type of a pull request (if it is Close, Commit, DESCR, Merge, PC, RC)</li> <li><em>ID</em> : specific ID of the comment</li> <li><em>Body</em> : contains the body of the comment analyzed</li> <li><em>Url</em> : link at each comments</li> <li><em>NumberPR</em> : number of a PR</li> <li><em>Name_Contributor</em>: contains the name of a contributor for each comment</li> <li><em>Contributor_classification</em>: contains the assignment of a specific classification (academic, bot, ML expert, software engineer, unknown), obtained after manual analysis, for each contributor of a comment</li> </ul> </li> <li> <p><code>Contributors_ManualAnalysis.csv</code>: contains the manual analysis performed by two authors to assign a classification for each contributor. The file contains the following columns:</p> <ul> <li><em>Contributors</em>: contains the name of the contributor for each comment</li> <li><em>Link GitHub</em>: contains the link to the GitHub page for each contributor</li> <li><em># PR</em>: contains the number of PRs in which a specific contributor is involved</li> <li><em>Annotator1</em>: manual classification of the first annotator</li> <li><em>Annotator2</em>: manual classification of the second annotator</li> <li><em>Final Classification</em>: contains the final label (academic, bot, ML expert, software engineer, unknown)assigned for each contributor</li> <li><em>Organization</em>: contains the organization, if any (Google, Hugging Face, Microsoft, OpenAI)</li> </ul> </li> <li> <p><code>ManualAnalysis.csv</code>: contains the manual analysis performed regarding Comment Type, Nature of Comment and Artifact. The .csv contains the following columns:</p> <ul> <li><em>URL</em>: contains the link to each comment</li> <li><em>NumberPR</em> : number of pull request</li> <li><em>CommentType1</em>: contains the classification of the first annotator in merit of comment type (Conventional review, Initial implementation, Management, ML review, Other)</li> <li><em>CommentType2</em>: contains the classification of the second annotator in merit of comment type (Conventional review, Initial implementation, Management, ML review, Other)</li> <li><em>NatureComment1</em>: contains the classification of the first annotator about the nature of the comment (Approval, Bug fix, Buid error, Clarification, Code, Code convention and spacing, Code review, Comment, Enhancement request, Explanation, Feedback, Introducing alternative implementation, Pinging, Plan for merging into TF, Question, References and referrals, Request a review, Request documentation improvement, Request test, Request verification, Review, Review assignment)</li> <li><em>NatureComment2</em>: contains the classification of the second annotator about the nature of the comment (Approval, Bug fix, Buid error, Clarification, Code, Code convention and spacing, Code review, Comment, Enhancement request, Explanation, Feedback, Introducing alternative implementation, Pinging, Plan for merging into TF, Question, References and referrals, Request a review, Request documentation improvement, Request test, Request verification, Review, Review assignment)</li> <li><em>Artifact1</em>: contains the classification of the first annotator with respect to the artifact (Article, Code, Commit, Issue/bug, Link, Review, Other)</li> <li><em>Artifact2</em>: contains the classification of the second annotator with respect to the artifact (Article, Code, Commit, Issue/bug, Link, Review, Other)</li> <li><em>FINALCommentType</em>: contains the final classification of the comment type after the resolution of the conflicts</li> <li><em>FINALNatureComment</em>: contains the final classification of the nature of the comment after resolution of conflicts</li> <li><em>FINALArtifact</em>: contains the final classification of the artifact after the resolution of conflicts</li> </ul> </li> <li> <p><code>Summary_PR.csv</code>: contains the details about the composition of each PRs. The file contains the following columns:</p> <ul> <li><em>#PullRequest</em>: contains the number of all pull requests analyzed</li> <li><em>#events</em>: contains the number of all the events analyzed for each PR</li> <li><em>#comments</em>: contains the number of all comments for each PR</li> <li><em>#Commit</em>:contains the number of Commit for each PR</li> <li><em>PC</em>: contains the number of PC for each PR</li> <li><em>RC</em>: contains the number of RC for each PR</li> </ul> </li> <li> <p><code>Total_PR_Comments.csv</code>: contains information about all comments analyzed. The columns are:</p> <ul> <li><em>Date</em>: contains the date of each comment</li> <li><em>Type</em>: describes the type of a pull request (if it is Close, Commit, DESCR, Merge, PC, RC)</li> <li><em>ID</em>: SHAn of the comment</li> <li><em>NumberPR</em>: PR number</li> <li><em>Name_Contributor</em>: contains the name of a contributor for each comment</li> <li><em>Body</em>: contains the body of the comment analyzed</li> <li><em>Url</em>: link at each comment</li> </ul> </li> </ul> <p>The replication also contains a directory <code>results</code> in which there are quantitative results. The directory contains:</p> <ul> <li> <p><code>Artifacts.csv</code>: This file contains the results for the artifact. The columns are:</p> <ul> <li><em>#PullRequest</em>: number of the pull request analyzed</li> <li><em>Article</em>: contains the percentage of the occurrences of the article in the specific pull request</li> <li><em>Code</em>: contains the percentage of the occurrences of the code in the specific pull request</li> <li><em>Commit</em>: contains the percentage of the occurrences of the commit in the specific pull request</li> <li><em>Issue reference</em>: contains the percentage of the occurrences of the issue reference in the specific pull request</li> <li><em>External link</em>: contains the percentage of the occurrences of the external link in the specific pull request</li> <li><em>Review</em>: contains the percentage of the occurrences of the review in the specific pull request</li> <li><em>Other</em>: contains the percentage of the occurrences of the other in the specific pull request</li> </ul> <p>&szlig;Also, the <code>Mean_Value</code> row contains the mean value of all occurrences for each column (Article, Code, Commit, Issue reference, External link, Review, Other)</p> </li> <li> <p><code>CommentType.csv</code>: this file contains the results for the comment type. The columns are:</p> <ul> <li><em>#PullRequest</em>:number of the pull request analyzed</li> <li><em>Conventional review</em>: contains the percentage of the occurrences of the conventional review in the specific pull request</li> <li><em>Initial implementation</em>: contains the percentage of the occurrences of the initial implementation in the specific pull request</li> <li><em>Management</em>: contains the percentage of the occurrences of the management in the specific pull request</li> <li><em>ML review</em>: contains the percentage of the occurrences of the ML review in the specific pull request</li> <li><em>Other</em>: contains the percentage of the occurrences of the Other in the specific pull request</li> </ul> <p>Also, the <code>Mean_Value</code> row contains the mean value of all occurrences for each column (Conventional review, Initial implementation, Management, ML review, Other)</p> </li> <li> <p><code>Contributor.csv</code>: this file contains the results for the contributors. The columns are:</p> <ul> <li><em>#PullRequest</em>: number of the pull requests analyzed</li> <li><em>academic</em>: contains the percentage of the occurrences of the academic in the specific pull request</li> <li><em>bot</em>: contains the percentage of the occurrences of the bot in the specific pull request</li> <li><em>ML expert</em>: contains the percentage of the occurrences of the ML expert in the specific pull request</li> <li><em>software engineer</em>: contains the percentage of the occurrences of the software engineer in the specific pull request</li> <li><em>unknown</em>: contains the percentage of the occurrences of unknown in the specific pull request</li> </ul> <p>Also, the <code>Mean_Value</code> row contains the mean value of all occurrences for each column (academic, bot, ML expert, software engineer, unknown)</p> </li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Replication Package of "Developer-Centric Code Readability Assessment: Are We There Yet?"

<p>This repository contains the datasets and the scripts to replicate our work "Developer-Centric Code Readability Assessment: Are We There Yet?".</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Performance regression testing initiatives: A systematic mapping - Replication Package

<p>This repository contains the supporting material for the manuscript titled "Performance regression testing initiatives: A systematic mapping", currently under revision to the Journal of Information and Software Technology.&nbsp;</p> <p>Repository Content:</p> <ul> <li> <p><strong>Protocol Performance Regression Testing.xlsx</strong>: it includes the *protocol* specification, i.e., the data extraction process used for the selection of the studies and gathering all relevant information. The outcome of the systematic mapping consists of 68 *selected studies* and information supporting the answers to four research questions: *RQ1*, *RQ2*, *RQ3*, and *RQ4*;</p> </li> <li> <p><strong>scripts.zip</strong>: it includes all the scripts used in the forward snowballing. Two additional files support the replication of the scripts: (1) README.md explains how to run the scripts; (2) requirements.txt file reports the installation prerequisites and their dependencies;</p> </li> <li> <p><strong>SM-PRT-overview.xlsx</strong>: this spreadsheet reports the information about tools and replication packages, benchmark suites and subject systems adopted per study;</p> </li> <li> <p><strong>README-SM-PRT-overview.md.docx</strong>: this file describes the columns of the spreadsheet above to facilitate the usage of the collected information.</p> </li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo32/100

[Replication Package] Why Personalizing Deep Learning-Based Code Completion Tools Matters

<p>This repository contains scripts, datasets and results of the work: <em>Why Personalizing Deep Learning-Based Code Completion Tools Matters</em></p> <p><strong>Note: you can access our scripts also in the official GitHub repository: <a href="https://github.com/Devy99/comp-personalization">https://github.com/Devy99/comp-personalization</a></strong></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Replication package of the Paper "On the Relationships between the Initial Ecology Indicators of OSS Projects and Their Long-Term Popularity: An Exploratory Study on GitHub"

<p>The dataset is collected from GitHub API and GitHub GHTorrent dataset. A brief description of each folder is provided below:</p> <p><strong>1. "Dataset_and_Code" folder</strong></p> <p>Contains the final dataset and algorithms</p> <p><strong>2. "Test_parameters" folders&nbsp;</strong></p> <p>Includes datasets under different parameters and the corresponding reproduction code, which corresponds to the first experiment of RQ1</p> <p><strong>3. "Compare_baseline"folder</strong></p> <p>Includes the dataset used by our method, the dataset used by the baseline method, and the reproduction code, corresponding to the second experiment of RQ1</p> <p><strong>4. "PLS" folder</strong></p> <p>Includes the dataset used by PLS and the corresponding reproduction code, which corresponds to experiment of RQ2</p> <p><strong>5. " Indicator_Calculation " folder</strong><br>It contains the calculation methods for various metrics in the paper, as well as the corresponding key&nbsp;files.&nbsp;</p> <p><span><strong>6. " Appendix " folder</strong><br></span><span>It includes supplementary materials such as the methodology for metric calculations to address the reviewers' questions.</span>&nbsp;</p> <p><strong>Note</strong></p> <p>We have provided a corresponding README file in each folder to help others reproduce our results</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Replication Package for "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models"

<p>This repository contains the replication package for the paper titled "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models". The README file and accompanying scripts provide detailed instructions to facilitate the replication of the analysis.</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record