Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,363
datasets available to search
ShareScore release 0.7.1
Dataset results
3,363 results for “Replication”
On-the-Fly Syntax Highlighting Using Neural Networks - Replication Package (Data)
<p>This dataset includes the data to replicate the study for the paper <em>On-the-Fly Syntax Highlighting Using Neural Networks</em>. It can be reused for future research in the field. We also include the detailed results obtained by executing our approach.</p> <p>HLNN-Resources.zip includes the input data already formatted to be directly used with the shared source code.</p> <p>The paper is published in the proceeding of the <em>30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE)</em>.</p>
Replication data for: Reconciliation k-median: Clustering with non-polarized representatives
<p># Description<br> These files contain the data employed in the experiments described in Bruno Ordozgoiti and Aristides Gionis. 2019. Reconciliation k-median: Clustering with Non-Polarized Representatives. In Proceedings of the 2019 World Wide Web Conference (WWW’19), May 13–17, 2019, San Francisco, CA, USA.</p> <p>Twitter ID's have been anonymized.</p> <p># Contents<br> domain_mentions.txt: Each line contains a domain name, a user ID and the number of times this user has mentioned this domain name in a tweet.<br> format: domain_name <TAB> user_id <TAB> mention_count</p> <p>domains_ideology_score.txt: Domain names and their ideology score, estimated as described in (Lahoti et al. WSDM 2018). Note: missing scores can be retrieved from supplementary data in https://doi.org/10.1093/poq/nfw006<br> format: domain_name <TAB> ideology_score</p> <p>follow_graph.txt: The Twitter follower graph. Each line contains a user id and the user id of one of its followers.<br> format: user_id <TAB> follower_user_id</p> <p>representatives.txt: US Congress representatives, each with Twitter handle and polarity score computed using Barbera's method (Barbera, 2015).<br> format: rep_name <TAB> website_url <TAB> district <TAB> twitter_handle <TAB> party <TAB> barbera_polarity_score</p> <p>user_polarity.txt: User ID's and polarity score computed using Barbera's method (Barbera, 2015).<br> format: user_id <TAB> barbera_polarity_score</p>
Replication Data for: Replication for: How Much Do Startups Impact Employment Growth in the U.S.?
<p>These are data files to support the replication of the blog post "How Much Do Startups Impact Employment Growth in the U.S.?" The replication is not a complete replication attempt. Files were downloaded from the U.S. Census Bureau at https://www.census.gov/ces/dataproducts/bds/data_firm.html on 2019-04-23. A codebook, as provided by the U.S. Census Bureau on the same date, is provided.</p>
Replication Data for: "Copularity of French and Dutch (semi-)copular constructions: a behavioral profile analysis"
<p>This data package contains all the data relevant to reproduce the results presented in the publication "Copularity of French and Dutch (semi-)copular constructions: a behavioral profile analysis".</p>
Bird Abundances at the Hubbard Brook Experimental Forest (1969-present) and on three replicate plots (1986-2000) in the White Mountain National Forest
Bird abundances have been determined from timed censuses, territory maps and nest locations at the Hubbard Brook Experimental Forest from 1969 to the present. This data set includes counts of the number of adult birds (males and females) per 10 ha at HBEF (1969 - present) and on three additional plots within the White Mountain National Forest (1986 - 2000). These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
Replication package of "Search-based Crash Reproduction using Behavioral Model Seeding"
<p>Search-based crash reproduction approaches assist developers during debugging by generating a test case which reproduces a crash given its stack trace. One of the fundamental steps of this approach is creating objects needed to trigger the crash. One way to overcome this limitation is seeding: using information about the application during the search process. With seeding, the existing usages of classes can be used in the<br> search process to produce realistic sequences of method calls which create the required objects. In this study, we introduce behavioral model seeding: a new seeding method which learns class usages from both<br> the system under test and existing test cases. Learned usages are then synthesized in a behavioral model (state machine). Then, this model serves to guide the evolutionary process. To assess behavioral model-seeding, we evaluate it against test-seeding (the state-of-the-art technique for seeding realistic objects) and no-seeding (without seeding any class usage). For this evaluation, we use a benchmark of 122 hard-to-reproduce crashes stemming from six open-source projects. Our results indicate that behavioral model-seeding outperforms both test seeding and no-seeding by a minimum of 6% without any notable negative impact on efficiency.</p>
# Replication code and data for: Tracking green space along streets of world cities
<p># Replication code and data for: Tracking green space along streets of world cities<br>Falchetta, G., & Hammad, A. T. (2025). Tracking green space along streets of world cities. Environmental Research: Infrastructure and Sustainability. https://doi.org/10.1088/2634-4505/add9c4 </p> <p>The file "gvi_358cities_2016_2023_yearly_falchetta_hammad.csv" contains<strong> output data</strong>, reporting sampling-point level data on the yearly (2016-2023) values of the Green View Index for the 190 cities covered in the paper AND an additional number of world cities (for a total of 358 cities). The "README_gvi_358cities_2016_2023_yearly_falchetta_hammad.txt" file contains a dictionary of each column name and units. </p> <p>____<br><br></p> <p>To replicate the analysis, the results, and the figures of the paper:</p> <ul> <li>Download input data from this Zenodo repository and code from Github https://github.com/giacfalk/urban_green_space_mapping_and_tracking</li> <li><em>*Optional data extraction steps* </em>(processed output data are already available in the Zenodo repository):<br> <ul> <li>Adjust your working directory</li> <li>Run [lines 4-11] of workflow/sourcer.R</li> <li>Run the Javascript scripts written by the string_generator_training.R and string_generator_prediction.R files in Google Earth Engine (https://code.earthengine.google.com) and complete the export to Drive tasks to generate the output .csv files</li> </ul> </li> <li>Run workflow/sourcer.R [lines 15-46] to train the ML model and make predictions (including figures and tables replication)</li> </ul> <div> <div> <div> </div> <div> <div> <div> </div> <div> <p dir="auto"> </p> <p dir="auto"> </p> </div> </div> </div> </div> </div> <div> <div> <div> </div> <div> <div> <div> </div> <div> <p dir="auto"> </p> <p dir="auto"> </p> </div> </div> </div> </div> </div>
Supplementary Datasets for the publication "Rousettus aegyptiacus Fruit Bats Do Not Support Productive Replication of Cedar Virus upon Experimental Challenge"
<p>Cedar henipavirus (CedV), which was isolated from the urine of pteropodid bats in Australia, belongs to the genus Henipavirus in the family of Paramyxoviridae. It is closely related to the Hendra virus (HeV) and Nipah virus (NiV), which have been classified at the highest biosafety level (BSL4) due to their high pathogenicity for humans. Meanwhile, CedV is apathogenic for humans and animals. As such, it is often used as a model virus for the highly pathogenic henipaviruses HeV and NiV. In this study, we challenged eight Rousettus aegyptiacus fruit bats of different age groups with CedV in order to assess their age-dependent susceptibility to a CedV infection. Upon intranasal inoculation, none of the animals developed clinical signs, and only trace amounts of viral RNA were detectable at 2 days post-inoculation in the upper respiratory tract and the kidney as well as in oral and anal swab samples. Continuous monitoring of the body temperature and locomotion activity of four animals, however, indicated minor alterations in the challenged animals, which would have remained unnoticed otherwise.</p>
Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing
<p><strong>OV2295 Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz: Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz: Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number </li> <li>minor_cn: HMM predicted minor copy number </li> <li>major_cn: HMM predicted major copy number </li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz: Table of SNVs per clone for OV2295 samples. Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het: is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format. Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: ‘SA922’: ‘OV2295(R2)’, ‘SA921’: ‘TOV2295(R)’, ‘SA1090’: ‘OV2295’,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>
Replication data for: "The hapax / type ratio: an indicator of minimally required sample size in productivity studies?"
<p>The dataset accompanies the scientific article "The hapax / type ratio: an indicator of minimally required sample size in productivity studies?" and can be used to reproduce the findings presented in this article. This dataset consists of two components, namely (i) the corpus data involving the Dutch semi-copular verb "raken" and (ii) an R analysis script to reproduce the computational steps.</p>
Replication data for: 'A First-Order Statistical Exploration of the Mathematical Limits of Micromagnetic Tomography'
<p>This repository contains the random data generated for obtaining results described in "A first-order statistical exploration of the mathematical limits of Micromagnetic Tomography". All tested parameters are systematically divided over different folders and subfolders. This dataset contains only .npy files, generated with python version 3.8.8 and numpy version 1.21.5.</p> <p>Each file can be opened with numpy.load(filename)</p> <p>The resulting figures are constructed with data of at least 15 iterations; each iteration is stored in a separate folder 'test_' followed by the iteration number.</p> <p>The README file inside provides a detailled overview of the files included.</p>
On the Effectiveness of Transfer Learning for Code Search - Replication Package
<p>This repository represents the replication package for the paper <em>On the Effectiveness of Transfer Learning for Code Search</em>.</p> <p>The paper is published in the journal <em>IEEE Transactions on Software Engineering (TSE)</em>.</p> <p>In this replication package, we provide all the data and scripts we used in our study.</p>
An External Replication on the Effects of Test-driven Development Using a Multi-site Blind Analysis Approach
<p>This dataset contains the <strong>unblinded </strong>version of the data collected and analyzed for the experiment reported in the paper. </p> <p>The semantics of the data can be found in the spreadsheet. For the formulas on how to obtain this data from the raw data, please see the paper. </p>
Replication Materials for Disclosure Limitation and Confidentality Protection in Linked Data
<p>These are the data and derived figures as used in the chapter by Abowd, Schmutte, and Vilhuber, "Disclosure Limitation and Confidentiality Protection in Linked Data"</p>
Replication material for paper "Freihardt (2025): Trapped by climate change? (In)voluntary immobility in Bangladesh. Regional Environmental Change. DOI 10.1007/s10113-025-02452-3."
<p>This is the data and replication code underlying the paper:</p> <p>Freihardt, J. Trapped by climate change? (In)voluntary immobility in Bangladesh. <em>Reg Environ Change</em> <strong>25</strong>, 117 (2025). https://doi.org/10.1007/s10113-025-02452-3</p>
Replication data for: Online Media Use and COVID-19 Vaccination in Real-World Personal Networks: Quantitative Study
<p>This is the replication data for the scientific paper titled "Online Media Use and COVID-19 Vaccination in Real-World Personal Networks: Quantitative Study" accepted for publication in the Journal of Medical Internet Research (JMIR). For details on how to use the data files, please consider the "supplementary_material.R" file or the "supplementary_material.pdf" where the variables of interest and R code are presented.</p> <p>For the code to run correctly, have the files "multilevel_labels.R" and "glm_labels.R" in the same working directory as the .R or .Rmd script. They are executed in the background, applying modifications to labels inside the regression tables. </p> <p> </p>
Replication data for "Climate change may induce connectivity loss and mountaintop extinction in Central American forests"
<p>Model code and predictor data underlying the publication "<strong>Climate change may induce connectivity loss and mountaintop extinction in Central American forests</strong>".</p>
Replication files for: Strongmen Cry Too: The Effect of Aerial Bombing on Voting for The Incumbent in Competitive Autocracies
<p>The NATO bombing of Yugoslavia, which lasted from March 24, 1999 until June 10, 1999, was the largest air campaign in Europe since the bombing of Britain and Germany in World War II. The air raids lasted for 78 days and hit 108 out of 160 municipalities, excluding Kosovo and Montenegro. The bombing was spread out and largely aimed at military barracks, industrial facilities, transportation networks, and communication lines. This repo provides a novel dataset with information on over 1,000 targets in the Federal Republic of Yugoslavia, including the date, location, target type, and fatalities. Included is also R code for the replication of my article "Strongmen Cry Too: The Effect of Aerial Bombing on Voting for The Incumbent in Competitive Autocracies" that was accepted for publication at Journal of Peace Research.</p>
Dataset and replication information for It's About Time: How to Study Intertemporal Choice in Systems Design
<p>Dataset and replication package for the paper <em>It's About Time: How to Study Intertemporal Choice in Systems Design</em> (Fagerholm, F., De los Ríos, A., Cárdenas Castro, C., Gil, J., Chatzigeorgiou, A., Ampatzoglou, A., Becker, C. (2023). It’s About Time: How To Study Intertemporal Choice in Systems Design. Information and Software Technology.). The dataset consists of answers to a scenario-based questionnaire that collects data on intertemporal choice in the context of software development. An analysis script is provided to show the details of the calculations and analyses performed for the paper. The replication package includes the protocol for data collection sessions and different versions of the task scenario and questionnaire. More information is given in the description file.</p>
Replication package for "Motivation in the Dynamics of European Youth Migration"
<p>Replication package for the paper "Motivation in the Dynamics of European Youth Migration". The package contains the data and the SPSS and Stata code for the the analyses presented in the paper. The original data from which the variables are extracted was collected within the EU Horizon 2020 project YMOBILITY (2015-2018).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.