Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

677

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

677 results for “Replication package”

Learn how ShareScore rates datasets ↗
zenodo32/100

Replication package for "Do LLMs Generate Code with Type Annotations?"

<div> <p>See the README.md file for more details.</p> </div>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Anonymous Replication Package

<p>Replication package for &quot;Handling Environmental Uncertainty in Design Time Access Control Analysis&quot;.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Conversational Agents in Software Engineering - Replication Package

<p>Replication Package for the research method conducted in the paper &quot;Conversational Agents in Software Engineering&quot;</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Replication package for: "Hobo Economicus"

<p>Leeson, Peter T., R. August Hardy, and Paola A. Suarez. &quot;Hobo Economicus.&quot; Economic Journal.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Replication package for: Measuring Unfair Inequality: Reconciling Equality of Opportunity and Freedom from Poverty

<p>Replication package for:</p> <p>Hufe, Paul, Kanbur, Ravi and Peichl, Andreas: Measuring Unfair Inequality: Reconciling Equality of Opportunity and Freedom from Poverty, Review of Economic Studies, fothcoming</p> <p>&nbsp;</p> <p>This replication package contains information/data/code to generate all tables, figures, and in-text figures of the paper.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Replication package for: "The Brides of Boko Haram: Economic Shocks, Marriage Practices, and Insurgency in Nigeria"

<p>Rexer, Jonah, &quot;The Brides of Boko Haram: Economic Shocks, Marriage Practices, and Insurgency in Nigeria&quot;,&nbsp;<em>Economic Journal</em>, forthcoming.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Replication package for: The Welfare Effects of Transportation Infrastructure Improvements

<p>This package contains all the code necessary to reproduce the figures and tables in Allen and Arkolakis (forthcoming) &quot;The Welfare Effects of Transportation Infrastructure Improvements.&quot; The Review of Economic Studies. Please use the larger (480.6 MB) zip file.</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

ARTE experiments replication package

<p>This dataset contains the replication package for the paper entitled &quot;ARTE: Automated Generation of Realistic Test Inputs for Web APIs&quot;</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Replication package for: Price Floors and Externality Correction

<p>Replication data for Griffith, O&#39;Connell and Smith (2022) &quot;Price Floors and Externality Correction&quot; Economic Journal</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Replication Package for Immigration and Redistribution

<p>Replication package for Immigration and Redistribution, MS27453.</p> <p>The replication package contains all the data--survey data and actual statistics on immigrants and non-immigrants--and scripts to replicate all the tables and figures in the paper. It also includes three files that explain how we computed the statistics and the sources we used.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Replication Package for 'Anchoring Code Understandability Evaluations Through Task Descriptions'

<p>Replication package for:</p> <p>&quot;Anchoring Code Understandability Evaluations Through Task Descriptions&quot;, to be presented at&nbsp;ICPC &#39;22: IEEE/ACM International Conference on Program Comprehension,&nbsp;May 21--22, 2022, Pittsburgh, Pennsylvania, United States</p> <ul> <li>&nbsp;The <em>data</em>&nbsp;folder contains dataset and Rmarkdown analysis script. We recommend loading the script with RStudio. For your convenience, we include a self-contained HTML outputted file.</li> <li>The <em>code snippets</em>&nbsp;folder contains the experimental code snippets and &nbsp;a representation with syntax highlighting in HTML. We have implemented the code in this way in LimeSurvey and think that for replication purposes it may be useful to include the markup in this replication package (see `LimeSurvey` subdirectory).</li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Replication Package for: Welfare and Redistribution in Residential Electricity Markets with Solar Power

<p>The Stata Do-Files, Matlab Scripts and Functions in this folder can be used to replicate the tables and graphs of the paper &quot;Welfare and Redistribution in Residential Electricity Markets with Solar Power&quot; by Fabian Feger, Nicola Pavanini, and Doina Radulescu.</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

Replication package for: Using Bid Rotation and Incumbency to Detect Collusion: A Regression Discontinuity Approach

<p>Replication file for Using Bid Rotation and Incumbency to Detect Collusion: A Regression Discontinuity Approach. The files create figures and tables reported in the text.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Replication Package for Immigration and Redistribution

<p>Replication package for Immigration and Redistribution by Alberto Alesina, Armando Miano, and Stefanie Stantcheva.</p> <p>The package contains all the survey data and the data on actual&nbsp;statistics about immigrants and non-immigrants, as well as the scripts needed to replicate the figures and tables in the paper. The package also contains three excel files detailing how we computed the actual statistics and the sources we used.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Replication Package (version 2) for "Information Constraints and Examination Quality in Patent Offices: The Effect of Initiation Lags" (International Journal of Industrial Organization, forthcoming)

<p>This replication package contains the datasets in Stata 16 format and the Stata do file to generate the results reported in &quot;Information Constraints and Examination Quality in Patent Offices: The Effect of Initiation Lags&quot; by Sadao Nagaoka and Isamu Yamauchi, to be published in International Journal of Industrial Organization.</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Replication Package for "Trust Enhancement Issues in Program Repair"

<p>This is the replication artifact for our work on &quot;Trust Enhancement Issues in Program Repair&quot;. The corresponding paper has been published at the International Conference of Software Engineering (ICSE) 2022, and is available under the following URL:&nbsp;<a href="https://doi.org/10.1145/3510003.3510040">https://doi.org/10.1145/3510003.3510040</a>. A pre-print of our work is available on arXiv:&nbsp;<a href="https://arxiv.org/pdf/2108.13064.pdf">https://arxiv.org/pdf/2108.13064.pdf</a>.</p> <p>The artifacts is organized in two parts:</p> <ol> <li>the artifacts for our <strong>developer survey</strong>, and</li> <li>the artifacts for our <strong>empirical assessment</strong> of state-of-the-art automated program repair (APR) techniques.</li> </ol> <p>&nbsp;</p> <p><strong>1. Survey Artifacts</strong></p> <p>The <code>survey</code>&nbsp;folder includes:</p> <ul> <li><code>Survey_Form.pdf</code> -- It shows the PDF version of the web form of our survey.</li> <li><code>Study_Results.pdf</code> -- It shows a summary of the questions and responses.</li> <li><code>Codebooks.xlsx</code> -- It shows all created codebooks.</li> <li><code>CodedResults.xlsx</code> -- It shows the responses for all questions, the corresponding coding, and statistics we applied during our analysis. Additionally, it includes plots for all responses and also the plots that are included in our paper.</li> </ul> <p>&nbsp;</p> <p><strong>2. Experiment Artifacts</strong></p> <p>The <code>experiments</code> folder includes:</p> <ul> <li><code>tools.md</code> -- It lists and describes the APR techniques that we used in our experiments.</li> <li><code>Results.xlsx</code> -- Contains the results of the experiments for each tool we considered, configuration details, and all the data from the ManyBugs benchmark.</li> <li><code>protocols/</code> -- This folder includes the analysis protocols, which describe for each tool &quot;how&quot; we extracted the values for our evaluation metrics (see Table 3 in our paper).</li> <li><code>results/</code> -- Contains the log files and relevant outputs for all tools and configurations. In particular, it includes the generated patches.</li> <li><code>subjects/</code> -- Contains the subjects taken from the <a href="https://repairbenchmarks.cs.umass.edu">ManyBugs</a> benchmark. For our experiments, we made some changes to the instrumentations, test-ids, etc. The file <code>meta-data.json</code> states the configurations, relevant test cases, etc.</li> <li><code>tool-snapshots/</code> -- Contains the snapshots for the tools, which we used in our evaluation.</li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Replication package for "Affordable Housing and City Welfare"

<p>The package contains all the code necessary to reproduce the figures and tables in Favilukis, Mabille and Van Nieuwerburgh (forthcoming), &quot;Affordable Housing and City Welfare.&quot; Review of Economic Studies.</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Geographic Diversity in Public Code Contributions — Replication Package

<p>Geographic Diversity in Public Code Contributions - Replication Package</p> <p>This document describes how to replicate the findings of the paper: Davide Rossi and Stefano Zacchiroli, 2022,&nbsp;<em>Geographic Diversity in Public Code Contributions - An Exploratory Large-Scale Study Over 50 Years</em>. In 19th International Conference on Mining Software Repositories (MSR &rsquo;22), May 23-24, Pittsburgh, PA, USA. ACM, New York, NY, USA, 5 pages.&nbsp;<a href="https://doi.org/10.1145/3524842.3528471">https://doi.org/10.1145/3524842.3528471</a></p> <p>This document comes with the software needed to mine and analyze the data presented in the paper.</p> <p>Prerequisites</p> <p>These instructions assume the use of the&nbsp;<a href="https://www.gnu.org/software/bash/">bash</a>&nbsp;shell, the&nbsp;<a href="https://www.python.org/">Python</a>&nbsp;programming language, the&nbsp;<a href="https://www.postgresql.org/">PosgreSQL</a>&nbsp;DBMS (version 11 or later), the&nbsp;<a href="https://facebook.github.io/zstd/">zstd</a>&nbsp;compression utility and various usual *nix shell utilities (cat, pv, &hellip;), all of which are available for multiple architectures and OSs.<br> It is advisable to create a&nbsp;<a href="https://docs.python.org/3/tutorial/venv.html">Python virtual environment</a>&nbsp;and install the following PyPI packages:</p> <pre><code>click==8.0.4 cycler==0.11.0 fonttools==4.31.2 kiwisolver==1.4.0 matplotlib==3.5.1 numpy==1.22.3 packaging==21.3 pandas==1.4.1 patsy==0.5.2 Pillow==9.0.1 pyparsing==3.0.7 python-dateutil==2.8.2 pytz==2022.1 scipy==1.8.0 six==1.16.0 statsmodels==0.13.2</code></pre> <p>Initial data</p> <ul> <li><code>swh-replica</code>, a PostgreSQL database containing a copy of Software Heritage data. The schema for the database is available at&nbsp;<a href="https://forge.softwareheritage.org/source/swh-storage/browse/master/swh/storage/sql/">https://forge.softwareheritage.org/source/swh-storage/browse/master/swh/storage/sql/</a>.<br> We retrieved these data from&nbsp;<a href="https://www.softwareheritage.org">Software Heritage</a>, in collaboration with the archive operators, taking an archive snapshot as of 2021-07-07. We cannot make these data available in full as part of the replication package due to both its volume and the presence in it of personal information such as user email addresses. However, equivalent data (stripped of email addresses) can be obtained from the Software Heritage archive dataset, as documented in the article: Antoine Pietri, Diomidis Spinellis, Stefano Zacchiroli,&nbsp;<em>The Software Heritage Graph Dataset: Public software development under one roof</em>. In proceedings of MSR 2019: The 16th International Conference on Mining Software Repositories, May 2019, Montreal, Canada. Pages 138-142, IEEE 2019.&nbsp;<a href="http://dx.doi.org/10.1109/MSR.2019.00030">http://dx.doi.org/10.1109/MSR.2019.00030</a>.<br> Once retrieved, the data can be loaded in PostgreSQL to populate&nbsp;<code>swh-replica</code>.</li> <li><code>names.tab</code>&nbsp;- forenames and surnames per country with their frequency</li> <li><code>zones.acc.tab</code>&nbsp;- countries/territories, timezones, population and world zones</li> <li><code>c_c.tab</code>&nbsp;- ccTDL entities - world zones matches</li> </ul> <p>Data preparation</p> <ul> <li> <p>Export data from the&nbsp;<code>swh-replica</code>&nbsp;database to create&nbsp;<code>commits.csv.zst</code>&nbsp;and&nbsp;<code>authors.csv.zst</code></p> <pre><code>sh&gt; ./export.sh</code></pre> </li> <li> <p>Run the authors cleanup script to create&nbsp;<code>authors--clean.csv.zst</code></p> <pre><code>sh&gt; ./cleanup.sh authors.csv.zst</code></pre> </li> <li> <p>Filter out implausible names and create&nbsp;<code>authors--plausible.csv.zst</code></p> <pre><code>sh&gt; pv authors--clean.csv.zst | unzstd | ./filter_names.py 2&gt; authors--plausible.csv.log | zstdmt &gt; authors--plausible.csv.zst</code></pre> </li> </ul> <p>Zone detection by email</p> <ul> <li> <p>Run the email detection script to create&nbsp;<code>author-country-by-email.tab.zst</code></p> <pre><code>sh&gt; pv authors--plausible.csv.zst | zstdcat | ./guess_country_by_email.py -f 3 2&gt; author-country-by-email.csv.log | zstdmt &gt; author-country-by-email.tab.zst</code></pre> </li> </ul> <p>Database creation and initial data ingestion</p> <ul> <li> <p>Create the PostgreSQL DB</p> <pre><code>sh&gt; createdb zones-commit</code></pre> <p>Notice that from now on when prepending the&nbsp;<code>psql&gt;</code>&nbsp;prompt we assume the execution of psql on the&nbsp;<code>zones-commit</code>&nbsp;database.</p> </li> <li> <p>Import data into PostgreSQL DB</p> <pre><code>sh&gt; ./import_data.sh</code></pre> </li> </ul> <p>Zone detection by name</p> <ul> <li> <p>Extract commits data from the DB and create&nbsp;<code>commits.tab</code>, that is used as input for the zone detection script</p> <pre><code>sh&gt; psql -f extract_commits.sql zones-commit</code></pre> </li> <li> <p>Run the world zone detection script to create&nbsp;<code>commit_zones.tab.zst</code></p> <pre><code>sh&gt; pv commits.tab | ./assign_world_zone.py -a -n names.tab -p zones.acc.tab -x -w 8 | zstdmt &gt; commit_zones.tab.zst</code></pre> Use&nbsp;<code>./assign_world_zone.py --help</code>&nbsp;if you are interested in changing the script parameters.</li> <li> <p>Ingest zones assignment data into the DB</p> <pre><code>psql&gt; \copy commit_zone from program 'zstdcat commit_zones.tab.zst | cut -f1,6 | grep -Ev ''\s$'''</code></pre> </li> </ul> <p>Extraction and graphs</p> <ul> <li> <p>Run the script to execute the queries to extract the data to plot from the DB. This creates&nbsp;<code>commit_zones_7120.tab</code>,&nbsp;<code>author_zones_7120_t5.tab</code>,&nbsp;<code>commit_zones_7120.grid</code>&nbsp;and&nbsp;<code>author_zones_7120_t5.grid</code>.<br> Edit&nbsp;<code>extract_data.sql</code>&nbsp;if you whish to modify extraction parameters (start/end year, sampling, &hellip;).</p> <pre><code>sh&gt; ./extract_data.sh</code></pre> </li> <li> <p>Run the script to create the graphs from all the previously extracted tabfiles.</p> <pre><code>sh&gt; ./create_stackedbar_chart.py -w 20 -s 1971 -f commit_zones_7120.grid -f author_zones_7120_t5.grid -o chart.pdf</code></pre> </li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Replication package for "Blended Modeling in Commercial and Open-source Model-Driven Software Engineering Tools: A Systematic Study"

<p>Replication package for the paper&nbsp;<em>Blended Modeling in Commercial and Open-source Model-Driven Software Engineering Tools: A Systematic Study</em>.</p> <p>Protocol</p> <ul> <li><code>/01-protocol/protocol.pdf</code></li> </ul> <p>Data &amp; analysis scripts</p> <p>This replication package is structured as follows:</p> <ul> <li><code>/02-search</code>&nbsp;- Detailed data on the&nbsp;<code>/academic</code>&nbsp;and&nbsp;<code>/grey literature</code>&nbsp;search.</li> <li><code>/03-tools</code>&nbsp;- Identified tools and inclusion/exclusion decisions.</li> <li><code>/04-classification_schema</code>&nbsp;- Classification framework and the corresponding data extraction form.</li> <li><code>/05-data</code>&nbsp;- Clean data in a processable form.</li> <li><code>/06-analysis</code>&nbsp;- Analysis scripts and results.</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Data Set and Replication Package of Paper on Handling Environmental Uncertainty in Design Time Access Control Analysis

<p>Data set and replication package for Paper &quot;Handling Environmental Uncertainty in Design Time Access Control Analysis&quot;.</p> <p>The data set contains an overview of used case studies, with illustrations and descriptions.</p> <p>The replication package contains the implemented application&nbsp;as well as model instances of every case study used for the evaluation.</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record