Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,363
datasets available to search
ShareScore release 0.7.1
Dataset results
3,363 results for “Replication”
Supplementary Materials and Replication Folder for Frank (2024)
<p>This zipped folder contains the data and code needed to replicate the main findings in</p> <p>E. G. Frank,<span> </span>The economic impacts of ecosystem disruptions: Costs from substituting biological pest control. <em>Science </em><strong>385</strong>, eadg0344 (2024). <a href="http://dx.doi.org/10.1126/science.adg0344">doi:10.1126/science.adg0344</a></p> <p>However, while the code to generate the health results is included, the heatlh data needs to be obtained separately through an application to the National Center for Health Statistics. See details here: https://www.cdc.gov/nchs/nvss/nvss-restricted-data.htm</p> <p> </p> <p> </p>
Replication Package for a Systematic Literature Mapping of Agility in Safety-Critical Software Development within the Aerospace Industry
<p>This file collection package facilitates the replication of a Systematic Literature Mapping (SLM) focused on Agility in Safety-Critical Software Development within the Aerospace Industry. Authored by J. Eduardo Ferreira Ribeiro, João Gabriel Silva, and Ademar Aguiar, this dataset is dedicated to improving transparency and reproducibility in this field of study and future research.</p> <p>Specifically, the package includes:</p> <ul> <li><a href="https://github.com/zemacedo99/Replication-Package-Builder/releases/tag/v1.0.2">Replication Package Builder Version 1.0.2</a></li> <li>A list of terms (both inclusion and exclusion) used to construct the research string.</li> <li>A list of venues unrelated to the research topic, to be excluded from the results.</li> <li>The inclusion and exclusion criteria applied during the study.</li> <li>Lists of publication results from indexing services like Scopus, IEEE Xplore, Science Direct, HAL Open Science, Springer Nature, and the ACM Digital Library are all provided in CSV file format.</li> <li>A list of all publications in CSV format, compiled after the automated exclusion phase using the established inclusion and exclusion criteria.</li> <li>Finally, a complete list of all publications, including those from Snowball sampling, in XLSX format was compiled after the manual exclusion phase using the established inclusion and exclusion criteria.</li> </ul> <p>Compiled and published on Saturday, September 14, 2024, this dataset is crucial for researchers seeking to replicate or extend the SLM's findings.</p> <p>Lastly, we thank J. Antonio Dantas Macedo for contributing to developing and providing this <a href="https://github.com/zemacedo99/Replication-Package-Builder">replication package builder</a>.</p>
Replication Package for "PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems"
<p>This package contains the traceback data, pre-trained models, and static word embeddings used in the paper, PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems.</p>
Replication Data for: Achieving Liquid Processors by Colloidal Suspensions for Reservoir Computing
<p>Overview</p> <p>This README provides instructions for accessing and managing the SQLite database digits_channel_A.db, digits_channel_B.db, and digits_channel_C.db which contains the recordings for all the channels A, B, and C of the digital oscilloscope. The database includes a table named digits_data with information about each recording, including metadata and audio data.</p> <p><br>Database Schema<br>Table: digits_data</p> <p> id: INTEGER PRIMARY KEY<br> A unique identifier for each record.</p> <p> person_id: INTEGER<br> Identifier for the person associated with the recorded digit.</p> <p> digit: INTEGER<br> The digit recorded (0-9).</p> <p> repeat: INTEGER<br> The repeat count of the recording.</p> <p> audio: BLOB<br> Binary Large Object to store audio data.</p> <p>Prerequisites</p> <p> Python 3.x<br> sqlite3 library (included with Python standard library)</p>
Replication Data For: Spiking patterns in the globus pallidus highlight convergent neural dynamics across diverse genetic dystonia syndromes
<div><strong>Human Globus Pallidum Single-Unit Activity Dataset in Genetic Dystonia Patients</strong></div> <div> </div> <div>This dataset consists of tabular data encompassing diverse neural features extracted from spiking trains of stable single-unit activity. These units were isolated from raw microelectrode recordings obtained from the globus pallidum of genetic dystonia patients who underwent globus pallidus internal (GPi) deep brain stimulation (DBS) surgery. The dataset includes anonymized patient IDs, details about the patient's genetic dystonia mutation, as well as information on the hemisphere and depth of microelectrode recordings (MER). Additionally, it features neural properties such as firing rate, spiking regularity, neural bursts, oscillations, and pause characteristics of isolated single-unit activities (SUAs).</div> <div> </div> <div>To process the raw MER, we applied a semi-parametric offline spike sorting algorithm to isolate SUAs. The SUAs were analyzed both in the temporal and frequency domains to derive a comprehensive set of features related to spiking patterns.</div> <div> </div> <div>For those interested in replicating or understanding the feature extraction process, the MATLAB source code is available in the <a href="github.com/ahmetofficial/Spike-Feature-Generator">Github repository</a>.</div>
Virophage replication mode drives ecological and evolutionary shifts in a host-virus-virophage system
<p>We studied how virophage replication modes affect the ecological dynamics of a host-virus-virophage system and the virophage’s evolutionary responses. By manipulating the level of virophage (Mavirus) integration into the host (Cafeteria burkhardae) alongside the Cafeteria roenbergensis virus (CroV), we found that higher integration improved host survival but decreased virophage reactivation. These communities had lower population densities and fewer fluctuations in host and virus populations, while virophage fluctuations increased. The virophage’s dual replication mode plays a key role in maintaining microbial community stability.</p>
Replication Package For An Extended Study of Syntactic Breaking Changes in the Wild
<p>This is the replication package associated with the paper titled 'An Extended Study of Syntactic Breaking Changes in the Wild' published under the Empirical Software Engineering journal.<br>Modern software applications rely heavily on the usage of libraries, which provide reusable functionality, to accelerate the development process. As libraries evolve and release new versions, the software systems that depend on those libraries (the clients) should update their dependencies to use these new versions as the new release could, for example, include critical fixes for security vulnerabilities. However, updating is not always a smooth process, as it can result in software failures in the clients if the new version<br>includes breaking changes. Yet, there is little research on how these breaking changes impact the client projects in the wild. <br>To identify if changes between two library versions cause breaking changes at the client end, we perform an empirical study on Java projects built using Maven. For the analysis, we used 18,415 Maven artifacts, which declared 142,355 direct dependencies, of which 71.60% were not up-to-date. We updated these dependencies and found<br>that 11.58% of the dependency updates contain breaking changes that impact the client. We further analyzed these changes in the library which impact the client projects and examine if libraries have adhered to the semantic versioning scheme when introducing breaking changes in their releases. Our results show that changes in transitive dependencies were a major factor in introducing breaking changes during dependency updates and almost half of the detected client impacting breaking changes violate the semantic versioning scheme by introducing breaking changes in non-Major upda</p>
Linked collectors and determiners for: eDNA‑based detection of the invasive crayfish Pacifastacus leniusculus in streams with a LAMP assay using dependent replicates to gain higher sensitivity.
Natural history specimen data linked to collectors and determiners held within, "eDNA‑based detection of the invasive crayfish Pacifastacus leniusculus in streams with a LAMP assay using dependent replicates to gain higher sensitivity". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63">https://bionomia.net/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63">https://gbif.org/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63</a>. Formatted as a Frictionless Data package.
Replication Data for the retroharmonize R Package Case Study: Working With Arab Barometer Surveys
<p>Replication datasets for the <a href="https://retroharmonize.dataobservatory.eu/articles/arabbarometer.html">retroharmonize Case Study: Working With Arab Barometer Surveys</a></p>
Replication Package for ICSE'21 paper - Representation of Developer Expertise in Open Source Software
<p>Replication package for ICSE'21 paper: Representation of Developer Expertise in Open Source Software.</p> <p>See README for details.</p>
Replication package for the paper "What do Developers Discuss about Code Comments"
<pre><code class="language-markdown"># RP-commenting-practices-multiple-sources Replication package for the paper "What do Developers Discuss about Code Comments?" ## Structure ``` Appendix.pdf Tags-topics.md Stack-exchange-query.md RQ1/ LDA_input/ combined-so-quora-mallet-metadata.csv topic-input.mallet LDA_output/ Mallet/ output_csv/ docs-in-topics.csv topic-words.csv topics-in-docs.csv topics-metadata.csv output_html/ all_topics.html Docs/ Topics/ RQ2/ datasource_rawdata/ quora.csv stackoverflow.csv manual_analysis_output/ stackoverflow_quora_taxonomy.xlsx ``` ## Contents of the Replication Package --- - **Appendix.pdf**- Appendix of the paper containing supplement tables - **Tags-topics.md** tags selected from Stack overflow and topics selected from Quora for the study (RQ1 & RQ2) - **Stack-exchange-query.md** the query interface used to extract the posts from stack exchnage explorer. - **RQ1/** - contains the data used to answer RQ1 - **LDA_input/** - input data used for LDA analysis - `combined-so-quora-mallet-metadata.csv` - Stack overflow and Quora questions used to perform LDA analysis - `topic-input.mallet` - input file to the mallet tool - **LDA_output/** - **Mallet/** - contains the LDA output generated by MALLET tool - **output_csv/** - `docs-in-topics.csv` - documents per topic - `topic-words.csv` - most relevant topic words - `topics-in-docs.csv` - topic probability per document - `topics-metadata.csv` - metadata per document and topic probability - **output_html/** - Browsable results of mallet output - `all_topics.html` - `Docs/` - `Topics/` - **RQ2/** - contains the data used to answer RQ2 - **datasource_rawdata/** - contains the raw data for each source - `quora.csv` - contains the processed dataset (like removing html tags). To know more about the preprocessing steps, please refer to the reproducibility section in the paper. The data is preprocessed using Makar tool. - `stackoverflow.csv` - contains the processed stackoverflow dataset. To know more about the preprocessing steps, please refer to the reproducibility section in the paper. The data is preprocessed using Makar tool. - **manual_analysis_output/** - `stackoverflow_quora_taxonomy.xlsx` - contains the classified dataset of stackoverflow and quora and description of taxonomy. - `Taxonomy` - contains the description of the first dimension and second dimension categories. Second dimension categories are further divided into levels, separated by `|` symbol. - `stackoverflow-posts` - the questions are labelled relevant or irrelevant and categorized into the first dimension and second dimension categories. - `quota-posts` - the questions are labelled relevant or irrelevant and categorized into the first dimension and second dimension categories. --- </code></pre> <p> </p>
Visual Tracking of Entire Bumblebee Colonies Using Novel Pipeline Finds No Evidence of Gut-Brain Axis (Replicates 1, 2)
<p>This archive contains raw data processed from video files taken of bumblebee colonies during replicates 1 and 2 of a study on the effect of the gut microbiome on social behaviour. Files ending with "_raw.csv" contain data on read tags, while ones ending with "_noID.csv" contain data on potential tags. Files are named as follows: R[replicate number][Baseline/Data][Day]R[recording session][HiveID][VideoID]</p>
Text-fig. 3. Phylogenetic relationship of Peignecyon felinoides n. gen. et n. sp., within some selected Amphicyonidae, and some extinct caniform carnivorans. Paramiacis exilis is the outgroup. Searches were performed by means of the Branch and Bound and a Bootstrap analysis through 1,000 replicates. One tree is obtained (length 73 steps, consistency index (CI) = 0.6301, retention index (RI) = 0.7000). The numbers below nodes are Bremer indices, and the numbers above nodes are Bootstrap support percentages (only shown ≥ 50). in A New Thaumastocyoninae (Amphicyonidae, Carnivora) From The Early Miocene Of Tuchořice, The Czech Republic
Text-fig. 3. Phylogenetic relationship of Peignecyon felinoides n. gen. et n. sp., within some selected Amphicyonidae, and some extinct caniform carnivorans. Paramiacis exilis is the outgroup. Searches were performed by means of the Branch and Bound and a Bootstrap analysis through 1,000 replicates. One tree is obtained (length 73 steps, consistency index (CI) = 0.6301, retention index (RI) = 0.7000). The numbers below nodes are Bremer indices, and the numbers above nodes are Bootstrap support percentages (only shown ≥ 50).
Escherichia coli DNA replication study: processed alignment data
<p>Genomes are replicated by large protein complexes called replisomes. In bacterial DNA replication, two replisomes replicate the DNA starting from the same origin site and proceeding in opposite directions. Understanding their movement in vivo has been challenging. We used quantitative genome sequencing to characterize the dynamics of bacterial replisomes at 5 different temperatures (17, 22, 27, 32 and 37 °C) in exponential growth (3 replicates) or in stationary phase (one experiment at 17, 27 and 37 °C).</p> <p>The data deposited here give the coordinates of the sequence reads (deposited under the BioProject PRJNA772106) covering the Escherichia coli str. K-12 substr. MG1655 complete genome (accession number U00096.3).</p> <p>The file archive contains data files for each sample, at nucleotide resolution and binned in intervals of 10,000 base pairs. It also contains a C program to perform the binning and a README summarising how the alignment was done. <em>Please note that once uncompressed, the data will take 5 Gb of disks space in total.</em></p>
Data for replication of the publication: Probabilistic leak localization in water distribution networks using a hybrid data-driven and model-based approach
<p>20 to 30% of drinking water produced is lost due to leaks in water distribution pipes. In times of water scarcity, losing so much treated water comes at a significant cost, both environmentally and economically. In this paper, we propose a hybrid leak localization approach combining both model-based and data-driven modeling. Pressure heads of leak scenarios are simulated using a hydraulic model, and then used to train a machine-learning based leak localization model. A key element of our approach is that discrepancies between simulated and measured pressures are accounted for using a dynamically calculated bias correction, based on historical pressure measurements. Data of in-field leak experiments in operational water distribution networks were produced to evaluate our approach on realistic test data. Two problematic settings for leak localization were examined. In the first setting, an uncalibrated hydraulic model was used. In the second setting, an extended version of the water distribution network was considered, where large parts of the network were insensitive to leaks. Our results show that the leak localization model is able to reduce the leak search region in parts of the network where leaks induce detectable drops in pressure. When this is not the case, the model still localizes the leak but is able to indicate a higher level of uncertainty with respect to its leak predictions.</p>
Replication Data for "How Closely are Common Mutation Operators Coupled to Real Faults?"
<p># Replication Data for "How Closely are Common Mutation Operators Coupled to Real Faults?"</p> <p>## Overview</p> <p>In mutation testing, faulty versions of a program are generated through automated modifications of source code. These mutants are used to assess and improve test suite quality, under the assumption that detection of mutants is indicative of a test suite's ability to detect real faults - i.e., that mutants and faults have a semantic relationship. Improving the effectiveness - in both cost and quality - of mutation testing may lie in better understanding this relationship, in particular with regard to how individual mutation operators (types) couple to real faults. </p> <p>In this study, we examine coupling between 32,002 mutants produced by 31 mutation operators and 144 real faults, using a scale based on number of failing tests and reasons for failure. Ultimately, we observed that 9.92% of the mutants are strongly coupled to real faults, and 51.03% of the faults have at least one strongly coupled mutant. We identify and examine mutation operators with the highest median coupling, as well as the operators that tend to produce non-compiling mutants, undetected mutants, and mutants that cause the most tests to fail outside of the tests that detect the actual fault. We also examine how coupling could be used to filter the set of operators employed, leading to potentially significant cost savings during mutation testing. Our findings could lead to improvements in how mutation testing is applied, improved implementation of specific mutation operators, and inspiration for new mutation operators. </p> <p>## Data Contained in This Package</p> <p>- mutant_data.csv</p> <p>This dataset contains the coupling results for all mutants considered in our experiments. It contains the following attributes for each mutant:</p> <p>-- Project name from Defects4J<br> -- Fault number from Defects4J<br> -- Mutation ID<br> -- Mutation operator<br> -- Number of trigger tests for the fault (tests that detect the real fault)<br> -- Number of failing test cases for the mutant (-1 indicates a compilation error)<br> -- The number of failing trigger tests for the mutant<br> -- The number of trigger tests that fail for the same reason the tests failed for the real fault.<br> -- The number of failing non-trigger tests.<br> -- The categorization of coupling. In order: Compile Error, Not Detected, No Substitution, Partial Test Substitution + Additional Tests Fail, Partial Test Substitution, Partial Substitution + Additional Tests Fail, Partial Substitution, Test Substitution + Additional Tests Fail, Test Substitution, Strong Substitution + Additional Tests Fail, Strong Substitution. </p> <p>- mutant_logs/{Project}/{Project}{Fault Number}output.txt</p> <p>The raw output log that resulted from executing test cases for each mutant for each case example used from Defects4J. Used to generate the dataset discussed above. Scripting for generating the dataset is also included.</p>
Replication Data for: Geometry-Complete Perceptron Networks for 3D Molecular Graphs
<p>Included are preprocessed data files for the Newtonian many-body systems modeling task described in our accompanying manuscript.</p>
Replication package for Workshop on Software Engineering 22' - What does the pytest plugins data say?
<p>This database stores the information used to run the experiment in the article: <strong>What does the pytest plugins data say?</strong></p>
Meta multiverse replication material
<p>Archived version (zip-file) of data, code and supplemental material for the publication "Meta-Analyzing the Multiverse: A Peek Under the Hood of Selective Reporting", published by Psychological Methods. A preprint is available at https://psyarxiv.com/43yae/ and the original (non-archived) material repository is available at https://osf.io/j8yg2/.</p>
De novo modelling of HEV replication polyprotein: Five-domain breakdown and involvement of flexibility in functional regulation - DATASET
<p>All models used for the paper "De novo modelling of HEV replication polyprotein: Five-domain breakdown and involvement of flexibility in functional regulation".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.