Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,363

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3,363 results for “Replication”

Learn how ShareScore rates datasets ↗
zenodo44/100

Replication package for "An Empirical Assessment of Best-Answer Prediction Models in Technical Q&A Sites" (EMSE 2018)

<p>Replication package for the paper:</p> <blockquote> <p>F. Calefato, F. Lanubile, and N. Novielli (2018) &ldquo;<a href="http://collab.di.uniba.it/fabio/wp-content/uploads/sites/5/2018/07/EMSE-D-17-00159_R3.compressed.pdf">An Empirical Assessment of Best-Answer Prediction Models in Technical Q&amp;A Sites</a>.&rdquo;&nbsp;Empirical Software Engineering Journal, DOI:&nbsp;<a href="http://dx.doi.org/10.1007/s10664-018-9642-5">10.1007/s10664-018-9642-5</a>.</p> </blockquote>

openother-openFeb 2019View details →
zenodo44/100

Compound annual growth rate for software: replication package

<p>This repository contains the reproducibility package (software and data) for the following paper.</p> <p>Les Hatton, Diomidis Spinellis, and Michiel van Genuchten. The long-term growth rate of evolving software: Empirical results and implications. <em>Journal of Software: Evolution and Process</em>, 29(5), May 2017. <a href="http://dx.doi.org/10.1002/smr.1847">doi:10.1002/smr.1847</a></p> <p>The amount of code in evolving software-intensive systems appears to be growing relentlessly, affecting products and entire businesses. Objective figures quantifying the software code growth rate bounds in systems over a large time scale can be used as a reliable predictive basis for the size of software assets. We analyze a reference base of over 404 million lines of open source and closed software systems to provide accurate bounds on source code growth rates. We find that software source code in systems doubles about every 42 months on average, corresponding to a median compound annual growth rate (CAGR) of 1.21&plusmn;0.01. Software product and development managers can use our findings to bound estimates, to assess the trustworthiness of road maps, to recognise unsustainable growth, to judge the health of a software development project, and to predict a system&rsquo;s hardware footprint.</p> <p>&nbsp;</p>

openapache2.0Jan 2017View details →
zenodo44/100

Dataset and replication package for Temporal Discounting in Software Engineering: A Replication Study

<p>Dataset and replication package for the paper Temporal Discounting in Software Engineering: A Replication Study (Fagerholm, F., Becker, C., Chatzigeorgiou, A., Betz, S., Duboc, L., Penzenstadler, B., Mohanani, R., Venters, C. (2019). Temporal Discounting in Software Engineering: A Replication Study. 13th ACM/IEEE International Symposium of Empirical Software Engineering and Measurement (ESEM 2019)). The dataset consists of answers to a questionnaire on temporal discounting in a technical debt context. Two questionnaire templates illustrate how to gather the data for professional and student participants. An analysis script is provided which shows the details of the calculations and analyses performed for the paper. More information is given in the description file.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Replication Data for: Determination of Intrinsic Effective Fields and Microwave Polarizations by High-Resolution Spectroscopy of Single NV Center Spins

<p>Data repository for: <strong>Determination of Intrinsic Effective Fields and Microwave Polarizations by High-Resolution Spectroscopy of Single NV Center Spins</strong></p> <p><em>Data description.pdf</em> describes the uploaded data.<br> <em>Data.xlsx</em> is the data represented in the paper.<br> <em>Esrfit_Npeak.m</em>, <em>Esrfit_xN.m</em>, <em>GaussianFunc.m</em>, <em>Gaussian_xN_Func.m</em>, <em>Lorentz_Func.m</em>, <em>Lorentz_xN_Func.m</em>, <em>Rabifit_xN.m</em>, <em>Rabi_xN_Func.m</em>, <em>FourierTransformRabi.m</em> are Matlab code files to transform and fit the data.</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Replication Package for the paper "Conversing with business process-aware Large Language Models: the BPLLM framework"

<p>Replication Package for the research paper "<em>Conversing with business process-aware Large Language Models: the BPLLM framework</em>".</p> <p>The package includes the process models, the questions (and expected answers), the results of the qualitative evaluation, and the Hugging Face links to the fine-tuned versions of Llama 3.1 8B employed in the quantitative evaluation of the framework.</p> <p>In particular, the process models are:</p> <ul> <li>The natural language Directly-follows graph (DFG) of the Food Delivery process: <em>food_delivery_activities.txt</em> for the definition of the activities and <em>food_delivery_flow.txt</em> for the sequence flow.</li> <li>The BPMN model of the Food Delivery, E-commerce, and Reimbursement processes: <em>ecommerce.bpmn</em>, <em>food_delivery.bpmn</em>, and <em>reimbursement.bpmn</em>.</li> </ul> <p>The datasets with the questions and the expected answers are:</p> <ul> <li><em>1_questions_answers_not_refined_for_DFG.csv</em> ;</li> <li><em>1.1_questions_answers_refined_for_DFG.csv</em> ;</li> <li><em>2_questions_answers_not_refined.csv</em> ;</li> <li><em>3_questions_answers_refined.csv</em> ;</li> <li><em>4_questions_answers_different_processes.csv</em> ;</li> <li><em>5_questions_answers_similar_processes.csv</em> ;</li> <li><em>6_questions_answers_refined_ft.csv</em> .</li> </ul> <p>The complete results of the qualitative evaluation are contained in the file <em>qualitative_experiments_results.pdf</em>.</p> <p>The Hugging Face links to the fine-tuned versions of Llama 3.1 8B are reported in <em>hf_links_finetuned_models.pdf</em>.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Replication package for paper: Insights on the Use of Software Design Principles in Machine Learning Pipelines

<p>This is the replication package of the paper "Insights on the Use of Software Design Principles in Machine Learning Pipelines".</p> <p>This replication package contains two files:</p> <ul> <li><a href="../api/records/13828806/draft/files/Data%20extraction.xlsx/content" target="_blank" rel="noopener noreferrer">Data extraction.xlsx</a>: file containing the details of the extracted data for each single ML project.&nbsp;</li> <li><a href="../api/records/13828806/draft/files/Source%20Code%20and%20Metadata.zip/content" target="_blank" rel="noopener noreferrer">Source Code and Metadata.zip</a>: zip file including the source code local copy analyzed and the repository metadata (.json) provided by GitHub API for each ML project .repository&nbsp;</li> </ul> <p>Reference: [1] Lidia L&oacute;pez, Cristina G&oacute;mez, and Claudia Ayala. Insights on the Use of Software Design Principles in Machine Learning Pipelines. <em>Accepted </em>in the 2024 edition of the International Conference on Product-Focused Software Process Improvement (PROFES 2024).</p> <p><strong>Note</strong>: The licence is applicable to the excel file. "Source Code and Metadata.zip" file contains source code repositories downloaded from GitHub, the license for each repository is defined in the corresponding GitHub repository by their authors.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Coat protein (CP) and trimmed replication-associated protein (Rep) amino acid alignments, phylogenetic analyses, and associated metadata for ICTV-approved begomovirus RefSeq species exemplars

<p>DATA RETRIEVAL</p> <p>Annotated begomovirus coding sequences corresponding to each begomovirus species exemplar with a RefSeq accession number listed in the ICTV Virus&nbsp;Metadata Resource (VMR #18, 2021-10-19,&nbsp;<a href="https://ictv.global/vmr">https://ictv.global/vmr</a>) were downloaded from GenBank in protein FASTA file format. CP and Rep amino acid sequences were extracted and split into separate data sets for analysis.&nbsp;We confirmed the identity of misannotated ORF&nbsp;products by performing a BLAST search.&nbsp;For exemplar sequences missing ORF annotations (listed in metadata spreadsheet), ORFfinder (<a href="https://www.ncbi.nlm.nih.gov/orffinder/">https://www.ncbi.nlm.nih.gov/orffinder/</a>) was used to identify CP and Rep ORFs that were subsequently translated and added to each corresponding data set after BLAST confirmation.</p> <p>ALIGNMENTS</p> <p>Multiple sequence alignments were constructed using the MUSCLE method (Edgar, 2004) as implemented in MEGA 11 (Tamura et al., 2021) and manually corrected using AliView v1.26<strong> </strong>(Larsson, 2014).&nbsp;After an initial alignment inspection, exemplars with either severely truncated (i.e., length &lt; 50% of the average length of the protein) or very divergent (i.e., causing us to doubt protein homology) CP or Rep sequences were excluded from the data set.&nbsp;Due to the difficulties in aligning the Rep sequences at the N- and C- terminal ends, the Rep alignment was trimmed to eliminate all residues prior to the iteron related domain (i.e., the known Rep functional region closest to the Rep start (Arguello-Astorga &amp; Ruiz-Medrano, 2001)) in the N-terminus and after a conserved geminivirus motif found near the C-terminus, which corresponds to where other circular, Rep-encoding single-stranded DNA viruses possess an arginine finger motif (Kazlauskas et al., 2019; Krupovic et al., 2020).&nbsp;In total, our CP and Rep data sets contained amino acid sequences from 432 begomovirus species exemplars that met our inclusion criteria.</p> <p>PHYLOGENETIC ANALYSIS</p> <p>Maximum likelihood (ML) trees were inferred with IQ-Tree v2.0.7 (Minh et al., 2020) using the best fitting substitution model identified by the built-in ModelFinder feature (Kalyaanamoorthy et al., 2017). Tree inference was performed with 3000 ultrafast bootstrap (UFBoot) replicates, a perturbation strength of 0.2 and a stopping rule requiring an iteration interval of 500 iterations between unsuccessful improvements to the local optimum. The -bnni flag was enabled to reduce the risk of overestimating branch supports with UFBoot due to severe model violations. The provided phylogenies in NEXUS format are midpoint-rooted and branches are colored based on traditional begomovirus geographic groupings:&nbsp;exemplars sampled in the Americas in orange and&nbsp;exemplars sampled in the &#39;Africa, Asia, Europe and Oceania&#39; (AAEO) region in blue.&nbsp;</p> <p>METADATA</p> <p>Metadata associated with each ICTV-approved species&nbsp;exemplar (n=445) &ndash; including country of isolation, geographic designation (i.e., AAEO/Americas), genome segmentation (i.e., monopartite/bipartite), presence/absence of V2/AV2 gene and length of genome/DNA-A segments &ndash; are included. Exemplars not incorporated into the other analyses&nbsp;are highlighted in red on the spreadsheet.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Replication Package for the Paper Titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction"

<p>This is a replication package for the paper titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction".</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Image-based Many-language Programming Language Identification - Replication Package

<p>This dataset contains the data, software, and instructions&nbsp;needed to replicate the findings of the paper:</p> <p>Francesca Del Bonifro, Maurizio Gabbrielli, Antonio Lategano, and Stefano Zacchiroli.&nbsp;Image-based Many-language<br> Programming Language Identification. <a href="https://peerj.com/computer-science/"><em>PeerJ Computer Science</em></a>, 2021 (to appear).&nbsp;DOI:&nbsp;<a href="https://dx.doi.org/10.7717/peerj-cs.631">10.7717/peerj-cs.631</a></p> <p>After retrieving the full dataset, extract the replication-package.zip archive&nbsp;and follow the instructions described in the README.md&nbsp;file.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Replication Package "Applying Test Case Prioritization to Software Microbenchmarks"

<p>Replication package for the paper &quot;Applying Test Case Prioritization to Software Microbenchmarks&quot; accepted for publication in Empirical Software Engineering.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Replication package for "Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation"

<p>This repository contains the replication package for the paper "Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation" by Fernando Vallecillos Ruiz, Anastasiia Grishina, Max Hort and Leon Moonen, accepted for publication in ACM Transactions on Software Engineering and Methodology on 2025-10-09.</p> <p>A preprint is deposited on arXiv with DOI: <a href="https://doi.org/10.48550/arXiv.2401.07994">10.48550/arXiv.2401.07994</a>.</p> <p>The replication package is archived on Zenodo with DOI: <a href="https://doi.org/10.5281/zenodo.10500593">10.5281/zenodo.10500593</a>.&nbsp;It is maintained on GitHub at <a href="https://github.com/secureIT-project/RTT_for_APR">https://github.com/secureIT-project/RTT_for_APR</a>.</p> <p>This project builds on code from the <a href="https://github.com/lin-tan/clm/">clm</a> project, which is (c) 2023, The ASSET research group led by Lin Tan,&nbsp;Purdue University, licensed under the BSD 3-Clause License (see jasper/LICENSE.BSD).&nbsp;All modifications and new contributions are (c) 2025 by the authors of this replication package&nbsp;and distributed under the MIT License (see LICENSE.MIT).&nbsp;The data, models and preprint are distributed under the CC BY 4.0 license.</p> <h2>Citation<code> </code></h2> <p>If you build on this data or code, please cite this work by referring to the paper:</p> <div> <pre><code>@article{ruiz2025:rtt, title = {Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation}, author = {Vallecillos Ruiz, Fernando and Anastasiia Grishina and Max Hort and Leon Moonen}, journal = {ACM Transactions on Software Engineering and Methodology (TOSEM)}, year = {2025}, publisher = {{ACM}} }</code></pre> </div> <h2>Organization</h2> <p>The replication package is organized as follows:</p> <ul> <li>clm-apr <ul> <li>plbart: code to generate patches with PLBART models.</li> <li>codet5: code to generate patches with CodeT5 models.</li> <li>transcoder: code to generate patches with the TransCoder model.</li> <li>incoder: code to generate patches with InCoder models.</li> <li>santacoder: code to generate patches with the SantaCoder model.</li> <li>starcoder: code to generate patches with the StarCoderBase model.</li> <li>quixbugs: code to validate patches generated for the QuixBugs benchmark.</li> <li>defects4j: code to validate patches generated for any of the Defects4J benchmarks.</li> <li>humaneval: code to validate patches generated for the HumanEval-Java benchmark.</li> </ul> </li> <li>humaneval-java: the HumanEval-Java benchmark proposed by Jiang et al. 2023</li> <li>jasper: a Java tool to parse Java programs needed to preprocess input.</li> <li>model: folder to download the language models.</li> <li>analysis_wandb: data from WandB and Jupyter notebook to create graphs.</li> <li>tmp_benchmarks: folder for temporary files used in patch validation. The folder may contain pairs of `paralell&rsquo; folders src and src_org for each benchmark, used to replace buggy code with candidate patches.</li> </ul> <h2>Replication</h2> <h3>Prerequisites</h3> <ul> <li>Python version: 3.8&mdash;3.10.</li> <li><a href="https://git-lfs.com/">Git LFS</a> is required for model downloading.</li> </ul> <h4>Weight and Biases (WandB)</h4> <ol> <li>Create an account on <a href="https://wandb.ai/">Weights and Biases</a></li> <li>Install the <a href="https://docs.wandb.ai/ref/python">Weights and Biases</a> library</li> <li>Run <code>wandb login</code> and follow the instructions</li> </ol> <h4>Set up OpenAI access</h4> <p>OpenAI account is needed with access to <code>gpt-3.5-turbo</code> and <code>gpt-4</code> . The <code>OPENAI_API_KEY</code> environment variable should be set to your OpenAI API access token.</p> <h3>Dependencies</h3> <ul> <li><a href="https://github.com/rjust/defects4j">Defects4J</a> - To generate inputs for the Defects4J datasets or to validate them, you need&nbsp;to have installed <a href="https://github.com/rjust/defects4j">their tool</a>.</li> <li>Java 8</li> <li>Apache Maven</li> </ul> <h3>Setup</h3> <p>We recommend the use of the setup script:</p> <pre><code>setup.sh </code></pre> <p>which performs the following:</p> <ol> <li>Creates a virtual environment for Python and activate it.</li> <li>Install the packages in <code>requirements.txt</code>.</li> <li>Compiles Jasper.</li> <li>Downloads parsers.</li> <li>Check if the Defects4J installation is correct.</li> </ol> <h3>Download models</h3> <p>The following bash script contains the code to download all of the models used:</p> <pre><code>models/download_models.sh </code></pre> <p>We recommend downloading only the models you are going to use due to their size</p> <pre><code>cd models chmod +x download_models.sh ./download_models.sh </code></pre> <p>To run one specific model, for example, PLBART (C#), use the following commands:</p> <pre><code>cd models git lfs install git clone https://huggingface.co/uclanlp/plbart-java-cs git clone https://huggingface.co/uclanlp/plbart-cs-java cd ../.. </code></pre> <h3>Step 1: Preprocessing and Prompting:</h3> <p>Each script in each <code>clm-apr/[model]</code> folder connects one or more models with<br>one dataset. These scripts follow the template: [benchmark]_[model]_[technique].py.<br>The scripts first create an <code>[model]_input.json</code> file with the preprocessed<br>input. Then generate outputs based on that file with one or more models.<br>For example:</p> <pre><code>cd clm-apr/plbart python quixbugs_plbart_round.py # Generates input for QuixBugs and generate patches using Java&lt;-&gt;C# RTT. python quixbugs_plbart_round_nl.py # Generates input for QuixBugs and generate patches using Java&lt;-&gt;NL RTT. </code></pre> <p>Optionally, use argument <code>--device_map cpu</code> if you wish to run the script on<br>CPU, for example:</p> <pre><code>python quixbugs_plbart_round.py --device_map cpu </code></pre> <p>Otherwise, the script will be run on all available CUDA GPU&rsquo;s.</p> <p>We have commented the generation of inputs in the scripts. Users are free to<br>uncomment this method and try for themselves. It is easily recognizable by<br>their name template <code>[model]_[benchmark]_input()</code>. In the previous case:</p> <pre><code>quixbugs_plbart_input() </code></pre> <h3>Step 2 and 3: Round Trip Translation and Postprocessing</h3> <p>These steps are also included in the [benchmark]_[model]_[technique].py<br>script mentioned above. They are modularized in the method recognizable by<br>their name template [model]_[benchmark]_output().<br>For example:</p> <pre><code>quixbugs_incoder_output() </code></pre> <p>This method:</p> <ol> <li>Reads the input json file.</li> <li>Generates outputs through the LLM.</li> <li>Postprocess the output (extract the patch, clean up extra token, etc.).</li> <li>Creates [model]_output_[technique]_[extra].json.</li> </ol> <p>The last 3 steps are repeated according to the number of runs set to performed<br>(10 in our experiments). Each run will produce a different file with the seed<br>used in its generation. For example, <code>quixbugs\_plbart\_round.py</code> and<br><code>quixbugs\_plbart\_round_nl.py</code> scripts create:</p> <pre><code>clm-apr/quixbugs/plbart_results/run_0/plbart_java_cs_java_output_round_csharp_batch.json clm-apr/quixbugs/plbart_results/run_0/plbart_java_nl_java_output_round_nl_batch.json </code></pre> <h3>Step 4: Evaluation of RTT Results:</h3> <p>The last step evaluates the generated outputs against the test-suites of each<br>benchmark. This script reads the previous outputs files and generates a new one<br>with the results of the test for one model. Furthermore, it connects with the<br><em>WandB</em> tool to calculate metrics and send them to analyze.</p> <p>Following the previous examples, to validate the results previously obtained,<br>we execute the following:</p> <pre><code>cd clm-apr/quixbugs python validate_quixbugs_parallel.py </code></pre> <p>Given the included JSON, this script would create:</p> <pre><code>clm-apr/quixbugs/plbart_results/run_0/plbart_java_cs_java_validate_round_csharp_batch.json </code></pre> <p>We have disabled <em>WandB</em> in the script to allow users to try the script first.<br>However, it can be easily activated by changing the parameter <code>mode="disabled"</code><br>to <code>mode="online"</code>.<br>We have set the variable <code>total_runs = 1</code>, as well as <code>input_file</code> and <code>output_file</code><br>to the results included. They should be modified accordingly to validate more runs<br>or to validate other files/models.</p> <h3>Included Results</h3> <p>We include two CSV files obtained through WandB.</p> <pre><code>'data_cleaned_grouped.csv': Aggregated metrics of the 25 outputs for all runs. 'full_data_all_runs.csv': All metrics for all outputs on all runs. </code></pre> <h2>Changelog</h2> <ul> <li>v1.0 - updates corresponding to the accepted version of the manuscript in TOSEM</li> <li>v0.1 - initial replication package corresponding to v1 of arXiv deposit: includes raw data, code, and example outputs.</li> </ul> <h2>References</h2> <p>Jiang, N.; Liu, K.; Lutellier, T.; and Tan, L. 2023. Impact of Code Language<br>Models on Automated Program Repair. In 45th International Conference on<br>Software Engineering (ICSE), 1430&ndash;1442. IEEE. ISBN 978-1-66545-701-9.</p> <div>&nbsp;</div>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Replication package for technical lag analysis for the JSEP journal.

<p>This is the replication package for our article &quot;A Formal Framework for Measuring Technical Lag in Component Repositories --- and its Application to npm&quot; submitted for the JSEP journal in 2018.</p> <p>This replication package requires Python 3.5+ to be installed, and all the dependencies listed in ``requirements.txt``.<br> They can be automatically installed using ``pip install -r requirements.txt``.&nbsp;<br> These experiment were executed on a Linux Ubuntu OS.</p> <p>To obtain the analysis used in the paper, one should execute ``jupyter notebook`` at the root of this replication package, and open the notebook contained in ``notebooks``.</p> <p>This replication package contains three folders (i.e scripts, notebooks and data), each folder has a README with a description of what it contains.</p> <p>The list of all npm package releases and Github repositories with their dependencies was download from the last available dataset of libraries.io: https://zenodo.org/record/1196312</p> <p>The data is under the Creative Commons Attribution Share-Alike 4.0 license.<br> The source code is under the GNU General Public License.</p> <p><br> For any more information about the details of these experiments, please contact: <strong><a href="mailto:ahmed.zerouali73@gmail.com">ahmed.zerouali73@gmail.com</a></strong></p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Codes to replicate statistical analysis in UPLIFT project Deliverable 2.4 Synthesis report, Chapter 6

<p>Policies attempting to mitigate the effects of urban inequality, often disregard affected citizens&rsquo; experiences, and thus fail to achieve&nbsp;maximum impact. By incorporating these perspectives into the policy design process, the project &quot;Urban PoLicy Innovation to address inequality with and for Future generaTions&quot;&nbsp;(UPLIFT), funded under the EU Horizon 2020 program&nbsp;aims to find innovative interventions in a bottom-up approach. The aims of UPLIFT project are to understand patterns and trends of inequality across Europe and to understand how individuals experience and adapt to inequality through participatory research. Moreover the project will together with the communities in four locations, co-design a policy tool aimed at addressing and reducing inequality and socio-economic divisions. The activity and results of the project can be followed at&nbsp;<a href="https://www.uplift-youth.eu/">https://www.uplift-youth.eu/</a>.</p> <p>Deliverable 2.4 (Synthesis report:&nbsp;socioeconomic inequalities in different urban contexts) is the final deliverable of work package 2 of the UPLIFT project, which aims to synthetize the main outcomes of the urban reports that described the policy environment around vulnerable individuals in the fields of education, employment and housing in 16 functional urban areas of the EU. In addition section 6.2 &quot;Statistical analysis of linkages between economic development of cities, their public policy performance and inequality outcomes&quot; of the report provides a statistical analysis of&nbsp;how local economic competitiveness&nbsp;and the local policy context affect urban deprivation and inequality among the young in European cities. The analysis is based on data from 2006, 2009, 2012, 2015 and 2019 of the Quality of Life in European Cities survey.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Replication package for: Solidarity and Fairness in Times of Crisis

<p>Replication package (code and data) for:</p> <blockquote> <p>Alexander W. Cappelen, Ranveig Falch, Erik &Oslash;. S&oslash;rensen and Bertil Tungodden (2021). Solidarity and fairness in times of crisis. Journal of Economic Behavior &amp; Organization 186: 1-11. <a href="https://doi.org/10.1016/j.jebo.2021.03.017">https://doi.org/10.1016/j.jebo.2021.03.017</a></p> </blockquote>

opencc-byDec 2022View details →
zenodo44/100

Studying Bug-Fixing Commits in the WoC Dataset: Replication Package

<p>A replication package for the MSR 2023 Challenge submission titled &quot;Studying Bug-Fixing Commits in the WoC Dataset&quot;.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Trade policy announcements can increase price volatility in global food commodity markets (Replication Data)

<p>Replication data for &quot;Trade policy announcements can increase price volatility in global food commodity markets&quot;:</p> <ul> <li>Original dataset on trade policy announcements from 2005 to 2017 for wheat and maize (corn) (details in codebook)</li> <li>Daily price ranges based on the highest and lowest price recorded for wheat and corn futures (traded at the Chicago Board of Trade, CBOT)</li> <li>Stocks-to-use data for the United States, which is compiled by the United States Department for Agriculture (USDA) and available at monthly frequency from their World Supply and Demand Estimates report</li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Data set for the replication package of the paper "Constriction of actin rings by passive crosslinkers"

<p>Data set for the replication package of the paper &quot;Constriction of actin rings by passive crosslinkers&quot;.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Electoral competition and strategic intra-coalition oversight in parliament: the case of the bipolar Belgian polity (replication data)

<p>Replication dataset for B. de Vet (2023). Electoral competition and strategic intra-coalition oversight in parliament: the case of the bipolar Belgian polity <em>(In: Political Studies Review)</em></p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

replicAnt - Plum2023 - 3D Models - Unreal Engine 5

<p>This dataset contains the 3D models used to generate all synthetic data presented in the&nbsp;<em>replicAnt -&nbsp;generating annotated images of animals in complex environments using Unreal Engine&nbsp;</em>manuscript. The models have been generated with the open-source photogrammetry platform <em>scAnt</em>&nbsp;<a href="https://peerj.com/articles/11155/">peerj.com/articles/11155</a>/ and&nbsp;have been pre-processed and converted into Unreal Engine 5 compatible .uasset files, to be used with the associated <em>replicAnt</em> project available from&nbsp;<a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created&nbsp;<em>replicAnt</em>, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. <em>replicAnt</em>&nbsp;places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with <em>replicAnt</em> can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, <em>replicAnt</em> may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College&rsquo;s President&rsquo;s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

replicAnt - Plum2023 - Detection & Tracking Datasets and Trained Networks

<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks used for detection and tracking experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data&nbsp;have been generated with the open-source photgrammetry platform scAnt&nbsp;<a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>.&nbsp;&nbsp;All synthetic data has been generated with the associated replicAnt project available from&nbsp;<a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created&nbsp;replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt&nbsp;places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Two video datasets were curated to quantify detection performance; one in laboratory and one in field conditions. The laboratory dataset consists of top-down recordings of foraging trails of <em>Atta vollenweideri</em> (Forel 1893) leaf-cutter ants. The colony was collected in Uruguay in 2014, and housed in a climate chamber at 25&deg;C and 60% humidity. A recording box was built from clear acrylic, and placed between the colony nest and a box external to the climate chamber, which functioned as feeding site. Bramble leaves were placed in the feeding area prior to each recording session, and ants had access to the recording area at will. The recorded area was 104 mm wide and 200 mm long. An OAK-D camera (OpenCV AI Kit: OAK-D, Luxonis Holding Corporation) was positioned centrally 195 mm above the ground. While keeping the camera position constant, lighting, exposure, and background conditions were varied to create recordings with variable appearance: The &ldquo;base&rdquo; case is an evenly lit and well exposed scene with scattered leaf fragments on an otherwise plain white backdrop. A &ldquo;bright&rdquo; and &ldquo;dark&rdquo; case are characterised by systematic over- or underexposure, respectively, which introduces motion blur, colour-clipped appendages, and extensive flickering and compression artefacts. In a separate well exposed recording, the clear acrylic backdrop was substituted with a printout of a highly textured forest ground to create a &ldquo;noisy&rdquo; case. Last, we decreased the camera distance to 100 mm at constant focal distance, effectively doubling the magnification, and yielding a &ldquo;close&rdquo; case, distinguished by out-of-focus workers. All recordings were captured at 25 frames per second (fps).<br> <br> The field datasets consists of video recordings of <em>Gnathamitermes</em>&nbsp;sp. desert termites, filmed close to the nest entrance in the desert of Maricopa County, Arizona, using a Nikon D850 and a Nikkor 18-105 mm lens on a tripod at camera distances between 20 cm to 40 cm. All video recordings were well exposed, and captured at 23.976 fps.<br> <br> Each video was trimmed to the first 1000 frames, and contains between 36 and 103 individuals. In total, 5000 and 1000 frames were hand-annotated for the laboratory-&nbsp;and field-dataset, respectively: each visible individual was assigned a constant size bounding box, with a centre coinciding approximately with the geometric centre of the thorax in top-down view. The size of the bounding boxes was chosen such that they were large enough to completely enclose the largest individuals, and was automatically adjusted near the image borders. A custom-written Blender Add-on aided hand-annotation: the Add-on is a semi-automated multi animal tracker, which leverages blender&rsquo;s internal contrast-based motion tracker, but also include track refinement options, and CSV export functionality. Comprehensive documentation of this tool and Jupyter notebooks for track visualisation and benchmarking is provided on the <a href="https://github.com/evo-biomech/replicAnt"><em>replicAnt</em></a>&nbsp;and <a href="https://github.com/FabianPlum/blenderMotionExport">BlenderMotionExport</a>&nbsp;GitHub repositories.</p> <p><strong>Synthetic data generation</strong></p> <p>Two synthetic datasets, each with a population size of 100, were generated from 3D models of \textit{Atta vollenweideri} leaf-cutter ants. All 3D models were created with the <em>scAnt</em>&nbsp;photogrammetry workflow. A &ldquo;group&rdquo; population was based on three distinct 3D models of an ant minor (1.1 mg), a media (9.8 mg), and a major (50.1 mg) (see <a href="https://zenodo.org/record/7849059">10.5281/zenodo.7849059</a>)). To approximately simulate the size distribution of <em>A. vollenweideri&nbsp;</em>colonies, these models make up 20%, 60%, and 20% of the simulated population, respectively. A 33% within-class scale variation, with default hue, contrast, and brightness subject material variation, was used. A &ldquo;single&rdquo; population was generated using the major model only, with 90% scale variation, but equal material variation settings.<br> <br> A <em>Gnathamitermes</em>&nbsp;sp. synthetic dataset was generated from two hand-sculpted models; a worker and a soldier made up 80% and 20% of the simulated population of 100 individuals, respectively with default hue, contrast, and brightness subject material variation. Both 3D models were created in Blender v3.1, using reference photographs.<br> <br> Each of the three synthetic datasets contains 10,000 images, rendered at a resolution of 1024 by 1024 px, using the default generator settings as documented in the Generator_example level file (see documentation on <a href="https://github.com/evo-biomech/replicAnt">GitHub</a>). To assess how the training dataset size affects performance, we trained networks on 100 (&ldquo;small&rdquo;), 1,000 (&ldquo;medium&rdquo;), and 10,000 (&ldquo;large&rdquo;) subsets of the &ldquo;group&rdquo; dataset. Generating 10,000 samples at the specified resolution took approximately 10 hours per dataset on a consumer-grade laptop (6 Core 4 GHz CPU, 16 GB RAM, RTX 2070 Super).</p> <p><br> Additionally, five datasets which contain both real and synthetic images were curated. These &ldquo;mixed&rdquo; datasets combine image samples from the synthetic &ldquo;group&rdquo; dataset with image samples from the real &ldquo;base&rdquo; case. The ratio between real and synthetic images across the five datasets varied between 10/1 to 1/100.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College&rsquo;s President&rsquo;s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record