Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

677

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

677 results for “replication package”

Learn how ShareScore rates datasets ↗
zenodo52/100

On-the-Fly Syntax Highlighting Using Neural Networks - Replication Package (Data)

<p>This dataset includes the data to replicate&nbsp;the study&nbsp;for the paper&nbsp;<em>On-the-Fly Syntax Highlighting Using Neural Networks</em>. It can be reused for future research in the field. We also include the detailed results obtained by executing our approach.</p> <p>HLNN-Resources.zip includes the input data already formatted to be directly used with the shared source code.</p> <p>The paper is published in the proceeding of the&nbsp;<em>30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE)</em>.</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

Replication package of "Search-based Crash Reproduction using Behavioral Model Seeding"

<p>Search-based crash reproduction approaches assist developers during debugging by generating a test case which reproduces a crash given its stack trace. One of the fundamental steps of this approach is creating objects needed to trigger the crash. One way to overcome this limitation is seeding: using information about the application during the search process. With seeding, the existing usages of classes can be used in the<br> search process to produce realistic sequences of method calls which create the required objects. In this study, we introduce behavioral model seeding: a new seeding method which learns class usages from both<br> the system under test and existing test cases. Learned usages are then synthesized in a behavioral model (state machine). Then, this model serves to guide the evolutionary process. To assess behavioral model-seeding, we evaluate it against test-seeding (the state-of-the-art technique for seeding realistic objects) and no-seeding (without seeding any class usage). For this evaluation, we use a benchmark of 122 hard-to-reproduce crashes stemming from six open-source projects. Our results indicate that behavioral model-seeding outperforms both test seeding and no-seeding by a minimum of 6% without any notable negative impact on efficiency.</p>

opencc-by-4.0Oct 2019View details →
zenodo48/100

On the Effectiveness of Transfer Learning for Code Search - Replication Package

<p>This repository represents the replication package for the paper <em>On the Effectiveness of Transfer Learning for Code Search</em>.</p> <p>The paper is published in&nbsp;the journal&nbsp;<em>IEEE Transactions on Software Engineering (TSE)</em>.</p> <p>In this replication package, we provide all the data and scripts we used in our study.</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

Replication package for "Motivation in the Dynamics of European Youth Migration"

<p>Replication package for the paper &quot;Motivation in the Dynamics of European Youth Migration&quot;. The package contains the data and the SPSS and Stata code for the&nbsp;the analyses presented in the paper. The original data from which the variables are extracted was collected within the EU Horizon 2020 project YMOBILITY (2015-2018).</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Measuring Software Testability Modulo Test Quality - Replication Package

<p>This repository represents the replication package for the paper&nbsp;<em>Measuring Software Testability Modulo Test Quality</em>.</p> <p>It includes the dataset and the Jupyter Notebook we used for the analysis in our paper.</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Replication package of "Good Things Come In Threes: Improving Search-based Crash Reproduction With Helper Objectives"

<p>The replication package for the study about using new helper objectives (MOHO) for crash reproduction. This study has been accepted at ASE 2020.</p> <p>&nbsp;</p> <p>Abstract:</p> <p>Evolutionary intelligence approaches have been successfully applied to assist developers during debugging by generating a test case reproducing reported crashes. These approaches use a single fitness function called&nbsp;<em>Crash Distance</em>&nbsp;to guide the search process toward reproducing a target crash. Despite the reported achievements, these approaches do not always successfully reproduce some crashes due to a lack of test diversity (premature convergence). In this study, we introduce a new approach, called&nbsp;<em>MO-HO</em>, that addresses this issue via multi-objectivization. In particular, we introduce two new Helper-Objectives for crash reproduction, namely&nbsp;<em>test length</em>&nbsp;(to minimize) and&nbsp;<em>method sequence diversity</em>&nbsp;(to maximize), in addition to&nbsp;<em>Crash Distance</em>.</p> <p>We assessed&nbsp;<em>MO-HO</em>&nbsp;using five multi-objective evolutionary algorithms (NSGA-II, SPEA2, PESA-II, MOEA/D, FEMO) on 124 hard-to-reproduce crashes stemming from open-source projects. Our results indicate that SPEA2 is the best-performing multi-objective algorithm for&nbsp;<em>MO-HO</em>.</p> <p>We evaluated this best-performing algorithm for&nbsp;<em>MO-HO</em>&nbsp;against the state-of-the-art: single-objective approach (Single-Objective Search) and decomposition-based multi-objectivization approach (<em>De-MO</em>). Our results show that&nbsp;<em>MO-HO</em>&nbsp;reproduces five crashes that cannot be reproduced by the current state-of-the-art. Besides,&nbsp;<em>MO-HO</em>&nbsp;improves the effectiveness (+10% and +8% in reproduction ratio) and the efficiency in 34.6% and 36% of crashes (i.e., significantly lower running time) compared to Single-Objective Search and&nbsp;<em>De-MO</em>, respectively. For some crashes, the improvements are very large, being up to +93.3% for reproduction ratio and -92% for the required running time.&nbsp;</p>

openother-openAug 2020View details →
zenodo44/100

The Daily Life of Software Engineers during the COVID-19 Pandemic -- Replication Package

<p>Following the onset of the COVID-19 pandemic and subsequent lockdowns, software engineers&#39; daily life was disrupted and abruptly forced into remote working from home. &nbsp;This change deeply impacted typical working routines, affecting both well-being and productivity.&nbsp;Moreover, this pandemic will have long-lasting effects in the software industry, with several tech companies allowing their employees to work from home indefinitely if they wish to do so. &nbsp;Therefore, it is crucial to analyze and understand how a typical working day looks like when working from home and how individual activities affect software developers&#39; well-being and productivity.&nbsp;We performed a two-wave longitudinal study involving almost 200 globally carefully selected software professionals, inferring daily activities with perceived well-being, productivity, and other relevant psychological and social variables.&nbsp;Results suggest that the time software engineers spent doing specific activities from home was similar when working in the office. (e.g., coding &gt; emails &gt; code review &gt; networking). &nbsp;However, we also found some meaningful mean differences.&nbsp;The amount of time developers spent on each activity was unrelated to their well-being, perceived productivity, and other variables.&nbsp;We conclude that working remotely is not per se&nbsp;a challenge for organizations or developers.</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Gender Differences in Public Code Contributions: a 50-year Perspective - Replication Package

<p>This page details the steps needed to replicate the findings of the paper:&nbsp;<a href="https://upsilon.cc/~zack/">Stefano Zacchiroli</a>,&nbsp;<em>Gender Differences in Public Code Contributions: a 50-year Perspective</em>,&nbsp;<a href="https://www.computer.org/csdl/magazine/so">IEEE Software</a>, 2021.</p> <p>After retrieving the replication package, follow the instruction described in the README.html file.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Replication package of https://doi.org/10.1088/1361-6595/abbae4

<p>This is the replication package of&nbsp;Plasma activation of N<sub>2</sub>, CH<sub>4</sub>&nbsp;and CO<sub>2</sub>: an assessment of the vibrational non-equilibrium time window, by A.W. van de Steeg et al.</p> <p>The zip file contains a general readme and all the required data on which the figures are based.&nbsp;</p> <p>Nearly all raw data is in .spe files, belonging to the Princeton Instruments camera used. The python library calibrate_fiber_pos can open the camera pictures.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches (Replication Package Part 3: OpenSSL dataset)

<h1><strong>The Replication Package of</strong></h1> <h1><strong>"Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches"</strong></h1> <h3><strong>Part 3 (OPENSSL Dataset)</strong></h3> <div> <div>This repository includes:</div> <ol> <li><em><strong>Code.zip</strong></em> that contains the codes to replicate some parts of this study:<br>a.&nbsp;<em>1_generate_datasets</em> implements our methodology to generate the datasets.<br>b.&nbsp;<em>2_run_models</em> runs the ML models during the evaluation.<br>c.&nbsp;<em>3_result_replication </em>generates charts presented in the paper from the ML evaluation results.</li> <li><em><strong>Datasets.zip</strong></em> that contain 2 folders:<br>a.&nbsp;<em>original</em> datasets: 1 from <a href="https://github.com/CGCL-codes/VulDeePecker" target="_blank" rel="noopener">NVD Vuldeepecker</a> and 3 extracted from&nbsp;<a href="https://github.com/ZeoVan/MSR_20_Code_vulnerability_CSV_Dataset" target="_blank" rel="noopener">BigVul</a>.<br> <div> <div>b. <em>OPENSSL</em> datasets: train, validation, test sets for each time of observation extracted using our methodology from <a href="https://github.com/ZeoVan/MSR_20_Code_vulnerability_CSV_Dataset" target="_blank" rel="noopener">BigVul</a>&nbsp;dataset for project <em>openssl</em>.</div> </div> </li> <li><em><strong>Pretrained-models.zip</strong></em>&nbsp;that we generated during our evaluation (3 test results for each time point in the timeline [2013-2019]).</li> <li><em><strong>Results.zip</strong></em> of our evaluation, the folder <em>ALL</em> contains the overall results and other folders are results by model.</li> </ol> <p><strong>UPDATED version 5<br></strong>- added a GLOBAL_README.md which contains the 3 stages and how they are connected to each other<br>- updated LineVul.ipynb: import AdamW from torch.optim instead of transformers<br>- updated README.md in Code2Vec with the prerequisites of Java to run gradlew for astmine</p> <p><strong>UPDATED version 6<br></strong>- updated CodeBert.ipynb: import AdamW from torch.optim instead of transformers</p> <p>Documentations</p> <ol> <li><em><strong>INSTALL.pdf&nbsp;</strong></em>: how to install the codes</li> <li><em><strong>README.pdf</strong></em>: readme file</li> <li><em><strong>REQUIREMENTS.pdf</strong></em>: hardware and software requirements</li> <li><em><strong>STATUS.pdf</strong></em>&nbsp;: status for artifact submission</li> <li><em><strong>LICENSE.pdf</strong></em>: the license of this artifact</li> <li><em><strong>PAPER.pdf</strong></em>: the camera-ready version of the paper</li> </ol> </div> <div> <div>Please refer to the following repositories for the other datasets and pre-trained models:</div> <div>- Part 1 NVD Vuldeeepecker :&nbsp;<a href="https://doi.org/10.5281/zenodo.8207883" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.8207883</a></div> - Part 2 LINUX :&nbsp;<a href="https://doi.org/10.5281/zenodo.10960662" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10960662</a><br> <div>- Part 4 POPPLER : <a href="https://doi.org/10.5281/zenodo.14713143">https://doi.org/10.5281/zenodo.14713143</a></div> <div>&nbsp;</div> <div>This work was partly funded by the EU under the H2020 Program AssureMOSS (Grant n. 952647) and the Horizon Europe Program Sec4AI4Sec (Grant n. 101120393), by the Italian Ministry of University and Research (MUR) under the P.N.R.R. &ndash; NextGenerationEU grant n.\ PE00000014 (SERICS subproject COVERT), and by the Dutch Research Council (NWO) under the grant NWA.1215.18.006 (Theseus) and grant KIC1.VE01.20.004 (HEWSTI).&nbsp;</div> </div>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Replication package and appendixes for Causal inference of server- and client-side code smells in web apps evolution

<p>-Analysis&nbsp;<br>--R scripts used to make the analisys, divided by folders<br>--Data folders used in the questions</p> <p>-Appendixes - used in the article to shwo extra tables and plots</p> <p>-data folders - Aggregation of data, each app has two files, CSV and xls</p> <p>-separated data folders - 5 files for each app, with lines corresponding to the each released official version<br>--serversmells<br>--clientsmells<br>--javascriptsmells<br>--Cloc(metrics)<br>--version (all oficial releases)</p> <p>-issues_bugs<br>--data -issues by app by release&nbsp;<br>--data_bugs_more - the same but only bugs, by app by release<br>--scripts - scrips used to aggregate issues (from daily issues to by release) anf the same for bugs</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Replication package for: "Is Secessionism Mostly About Income or Identity? A Global Analysis of 3,153 Subnational Regions"

<p>This repository contains the data and code to replicate the analyses performed in&nbsp;<a href="https://academic.oup.com/ej/article/135/668/1261/7918442?utm_source=authortollfreelink&amp;utm_campaign=ej&amp;utm_medium=email&amp;guestAccessKey=d6c8adb1-257c-47ad-827c-79b94cf86664" target="_blank" rel="noopener">"Is Secessionism Mostly About Income or Identity? A Global Analysis of 3,153 Subnational Regions"</a>&nbsp;by&nbsp;<a href="https://people.smu.edu/kdesmet/">Klaus </a><a href="https://people.smu.edu/kdesmet/" target="_blank" rel="noopener">Desmet</a>,&nbsp;<a href="https://sites.google.com/view/ignacioortuno" target="_blank" rel="noopener">Ignacio Ortu&ntilde;o-Ort&iacute;n</a>, and&nbsp;<a href="http://omerozak.com">&Ouml;mer </a><a href="http://omerozak.com" target="_blank" rel="noopener">&Ouml;zak</a>. If you use the code or data in this repository, please cite both the original paper and the dataset.<br><br>Citation:</p> <p>Desmet, Klaus, Ortu&ntilde;o-Ort&iacute;n, Ignacio, and &Ouml;zak, &Ouml;mer. (2024) "<a href="https://academic.oup.com/ej/article/135/668/1261/7918442?utm_source=authortollfreelink&amp;utm_campaign=ej&amp;utm_medium=email&amp;guestAccessKey=d6c8adb1-257c-47ad-827c-79b94cf86664" target="_blank" rel="noopener">Is Secessionism Mostly About Income or Identity? A Global Analysis of 3,153 Subnational Regions</a>", Economic Journal, Volume 135, Issue 668, May 2025, Pages 1261&ndash;1299.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Perceptions on the utility of community question and answer websites like Stack Overflow to software developers (Replication package)

<p>Interview Questions on the perception of the utility of CQAs like Stack Overflow to software developers. In this study, we focused on the questions highlighted in yellow.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Replication Package for the Paper Titled "How Well Do Software Practitioners Fix Code Vulnerabilities with Different Types of Explanations?"

<p>This is a replication package for the article 'How Well Do Software Practitioners Fix Code Vulnerabilities with Different Types of Explanations?'. The survey questions can be found here, and we encourage the survey to be re-used.</p> <p>We also include survey data (with demographic data and qualitative responses removed for anonymity reasons).</p> <p>The project team consists of Tracy Hall, Emily Winter, Fahad Al Debeyan (Lancaster University) and Lech Madeyski (Wroclaw University of Science and Technology). If you have any questions about the re-use of this survey, feel free to contact Fahad at&nbsp;<a href="mailto:e.winter@lancaster.ac.uk">f.aldebeyan@lancaster.ac.uk</a>.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Replication package for "An Empirical Assessment of Best-Answer Prediction Models in Technical Q&A Sites" (EMSE 2018)

<p>Replication package for the paper:</p> <blockquote> <p>F. Calefato, F. Lanubile, and N. Novielli (2018) &ldquo;<a href="http://collab.di.uniba.it/fabio/wp-content/uploads/sites/5/2018/07/EMSE-D-17-00159_R3.compressed.pdf">An Empirical Assessment of Best-Answer Prediction Models in Technical Q&amp;A Sites</a>.&rdquo;&nbsp;Empirical Software Engineering Journal, DOI:&nbsp;<a href="http://dx.doi.org/10.1007/s10664-018-9642-5">10.1007/s10664-018-9642-5</a>.</p> </blockquote>

openother-openFeb 2019View details →
zenodo44/100

Compound annual growth rate for software: replication package

<p>This repository contains the reproducibility package (software and data) for the following paper.</p> <p>Les Hatton, Diomidis Spinellis, and Michiel van Genuchten. The long-term growth rate of evolving software: Empirical results and implications. <em>Journal of Software: Evolution and Process</em>, 29(5), May 2017. <a href="http://dx.doi.org/10.1002/smr.1847">doi:10.1002/smr.1847</a></p> <p>The amount of code in evolving software-intensive systems appears to be growing relentlessly, affecting products and entire businesses. Objective figures quantifying the software code growth rate bounds in systems over a large time scale can be used as a reliable predictive basis for the size of software assets. We analyze a reference base of over 404 million lines of open source and closed software systems to provide accurate bounds on source code growth rates. We find that software source code in systems doubles about every 42 months on average, corresponding to a median compound annual growth rate (CAGR) of 1.21&plusmn;0.01. Software product and development managers can use our findings to bound estimates, to assess the trustworthiness of road maps, to recognise unsustainable growth, to judge the health of a software development project, and to predict a system&rsquo;s hardware footprint.</p> <p>&nbsp;</p>

openapache2.0Jan 2017View details →
zenodo44/100

Dataset and replication package for Temporal Discounting in Software Engineering: A Replication Study

<p>Dataset and replication package for the paper Temporal Discounting in Software Engineering: A Replication Study (Fagerholm, F., Becker, C., Chatzigeorgiou, A., Betz, S., Duboc, L., Penzenstadler, B., Mohanani, R., Venters, C. (2019). Temporal Discounting in Software Engineering: A Replication Study. 13th ACM/IEEE International Symposium of Empirical Software Engineering and Measurement (ESEM 2019)). The dataset consists of answers to a questionnaire on temporal discounting in a technical debt context. Two questionnaire templates illustrate how to gather the data for professional and student participants. An analysis script is provided which shows the details of the calculations and analyses performed for the paper. More information is given in the description file.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Replication Package for the paper "Conversing with business process-aware Large Language Models: the BPLLM framework"

<p>Replication Package for the research paper "<em>Conversing with business process-aware Large Language Models: the BPLLM framework</em>".</p> <p>The package includes the process models, the questions (and expected answers), the results of the qualitative evaluation, and the Hugging Face links to the fine-tuned versions of Llama 3.1 8B employed in the quantitative evaluation of the framework.</p> <p>In particular, the process models are:</p> <ul> <li>The natural language Directly-follows graph (DFG) of the Food Delivery process: <em>food_delivery_activities.txt</em> for the definition of the activities and <em>food_delivery_flow.txt</em> for the sequence flow.</li> <li>The BPMN model of the Food Delivery, E-commerce, and Reimbursement processes: <em>ecommerce.bpmn</em>, <em>food_delivery.bpmn</em>, and <em>reimbursement.bpmn</em>.</li> </ul> <p>The datasets with the questions and the expected answers are:</p> <ul> <li><em>1_questions_answers_not_refined_for_DFG.csv</em> ;</li> <li><em>1.1_questions_answers_refined_for_DFG.csv</em> ;</li> <li><em>2_questions_answers_not_refined.csv</em> ;</li> <li><em>3_questions_answers_refined.csv</em> ;</li> <li><em>4_questions_answers_different_processes.csv</em> ;</li> <li><em>5_questions_answers_similar_processes.csv</em> ;</li> <li><em>6_questions_answers_refined_ft.csv</em> .</li> </ul> <p>The complete results of the qualitative evaluation are contained in the file <em>qualitative_experiments_results.pdf</em>.</p> <p>The Hugging Face links to the fine-tuned versions of Llama 3.1 8B are reported in <em>hf_links_finetuned_models.pdf</em>.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Replication package for paper: Insights on the Use of Software Design Principles in Machine Learning Pipelines

<p>This is the replication package of the paper "Insights on the Use of Software Design Principles in Machine Learning Pipelines".</p> <p>This replication package contains two files:</p> <ul> <li><a href="../api/records/13828806/draft/files/Data%20extraction.xlsx/content" target="_blank" rel="noopener noreferrer">Data extraction.xlsx</a>: file containing the details of the extracted data for each single ML project.&nbsp;</li> <li><a href="../api/records/13828806/draft/files/Source%20Code%20and%20Metadata.zip/content" target="_blank" rel="noopener noreferrer">Source Code and Metadata.zip</a>: zip file including the source code local copy analyzed and the repository metadata (.json) provided by GitHub API for each ML project .repository&nbsp;</li> </ul> <p>Reference: [1] Lidia L&oacute;pez, Cristina G&oacute;mez, and Claudia Ayala. Insights on the Use of Software Design Principles in Machine Learning Pipelines. <em>Accepted </em>in the 2024 edition of the International Conference on Product-Focused Software Process Improvement (PROFES 2024).</p> <p><strong>Note</strong>: The licence is applicable to the excel file. "Source Code and Metadata.zip" file contains source code repositories downloaded from GitHub, the license for each repository is defined in the corresponding GitHub repository by their authors.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Replication Package for the Paper Titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction"

<p>This is a replication package for the paper titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction".</p>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record