Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
764
datasets available to search
ShareScore release 0.7.1
Dataset results
764 results for “Reproducibility”
Expanding the space of self-reproducing ribozymes using probabilistic generative models
<p>This repository contains the code and data produced in "Expanding the space of self-reproducing ribozymes using probabilistic generative models".</p>
Data and scripts to reproduce figures from "Ocean warming threatens the viability of 60% of Antarctic ice shelves"
<p>This is the formatted data related to the paper "Ocean warming threatens the viability of 60% of Antarctic ice shelves".</p> <p>Main data used to produce the main figures:<br>> bayesian_weights_davison_varying_combined_2300_withoutGISS.nc: Bayesian weights used in the computation of weighted likelihoods and averages considering all simulations going to 2300 - used in the main study<br>> all_fluxes_br_withoutGISS.nc: all fluxes forming the ice-shelf mass balance propagated across the different dimensions of uncertainty (CMIP models, basal melt parameterisations, bed plasticities)<br>> hydrofracturing_limits_new.nc: File containing the percentiles of reaching the hydrofracturing criterion</p> <p>Files needed to compute the weights:<br>> area_isf_greene22.nc: ice-shelf area used in Davison et al. 23 (from <a href="https://doi.org/10.1038/s41586-022-05037-w">Greene et al. 2022</a>)<br>> varying_conditions_davison23.nc: Ice-shelf mass budget fluxes from 1997 on from <a href="https://doi.org/10.1126/sciadv.adi0186">Davison et al. 2023</a><br>> steadystate_davison23.nc: Ice-shelf mass budget fluxes from the steady state from <a href="https://doi.org/10.1126/sciadv.adi0186">Davison et al. 2023</a></p> <p>Geometric information:<br>> gridarea_ISMIP6_AIS_4000m_grid.nc: File containing area of grid cells<br>> Mask_Iceshelf_4km_IMBIE_withNisf.nc: Mask used to define the location of the different ice shelves, based on <a href="https://zenodo.org/records/15863352">Caillet et al. 2025 </a><br>> BedMachine_4km_slope_info_bedrock_draft_latlon_oneFRIS.nc: Geometric information needed for neural network parameterisation, inferred using the <a href="https://github.com/ClimateClara/multimelt">multimelt package</a><br>> BedMachinev2_4km_isf_masks_and_info_and_distance_oneFRIS.nc: Geometric information needed for different computations, inferred using the <a href="https://github.com/ClimateClara/multimelt">multimelt package</a><br>> ano_choice_NEMOorISMIP_withoutGISS.nc: file recording which T and S profiles should be taken, either ISMIP climatology or NEMO hindcast, depending on their difference to the reference mass balance</p> <p>Data for analysis only until 2100:<br>> bayesian_weights_davison_varying_combined_withoutGISS.nc: Bayesian weights used in the computation of weighted likelihoods and averages considering all simulations going to 2100 (see Extended Data Fig. 5)</p> <p>Additional files needed for the additional analysis on other plausible geometries: <br>> ElmerIce_4km_2100isf_masks_and_info_and_distance_oneFRIS.nc: Geometric information needed for different computations for the 2100 ice-sheet geometry<br>> ElmerIce_4km_2100_slope_info_bedrock_draft_latlon_oneFRIS.nc: Geometric information needed for neural network parameterisation for the 2100 ice-sheet geometry<br>> ElmerIce_4km_2150isf_masks_and_info_and_distance_oneFRIS.nc: Geometric information needed for different computations for the 2150 ice-sheet geometry<br>> ElmerIce_4km_2150_slope_info_bedrock_draft_latlon_oneFRIS.nc: Geometric information needed for neural network parameterisation for the 2150 ice-sheet geometry<br>> all_fluxes_br_withoutGISS_ElmerIcegeo2100.nc: all fluxes forming the ice-shelf mass balance propagated across the different dimensions of uncertainty (CMIP models, basal melt parameterisations, bed plasticities) for the 2100 ice-sheet geometry<br>> all_fluxes_br_withoutGISS_ElmerIcegeo2150.nc: all fluxes forming the ice-shelf mass balance propagated across the different dimensions of uncertainty (CMIP models, basal melt parameterisations, bed plasticities) for the 2150 ice-sheet geometry<br><br>The scripts are explained in the associated README in scripts_paper_iceshelf_viability.zip. You can also find them here: https://github.com/ClimateClara/scripts_paper_iceshelf_viability<br><br>Here is a summary of the README:</p> <p>- Figure 2 and 3 were done with /notebooks_for_figures/timeseries_nb_viable_isf_withoutGISS_withhydrofrac_calving0.ipynb and /notebooks_for_figures/2D_subplots_viability_proba_withoutGISS_calving0.ipynb<br>- Figure 4 was done with /notebooks_for_figures/2D_subplots_viability_proba_withoutGISS_calving0.ipynb<br>- Figure 5 was done with /notebooks_for_figures/timeseries_nb_viable_isf_withoutGISS_withhydrofrac_calving0.ipynb and notebooks_for_figures/2D_subplots_viability_proba_onlyhydrofrac.ipynb</p> <p>- Extended Data Figures 1 to 4 were done with /notebooks_for_figures/plot_mass_fluxes.ipynb<br>- Extended Data Figure 5 was done with /notebooks_for_figures/timeseries_nb_viable_isf_withoutGISS_withhydrofrac_calving0.ipynb<br>- Extended Data Figure 6 was done with /notebooks_for_figures/timeseries_nb_viable_isf_withoutGISS_withhydrofrac.ipynb and /notebooks_for_figures/2D_subplots_viability_proba.ipynb<br>- Extended Data Figure 7 was done with /notebooks_for_figures/2D_subplots_viability_proba.ipynb<br>- Extended Data Figure 8 was done with /notebooks_for_figures/timeseries_nb_viable_isf_withoutGISS_withhydrofrac_ElmerIcegeometries.ipynb and /notebooks_for_figures/2D_subplots_viability_proba_withoutGISS_ElmerIcegeometries.ipynb<br>- Extended Data Figures 10 to 12 were done with /notebooks_for_figures/histo_weights_new.ipynb</p> <p>In the folder 'notebooks_for_datapreparation', you will find a few scripts to prepare the data. These are not as detailed and not meant to be run out of the box but permit to give insight into the practical application of the methods described in the paper. For potential inspiration of similar work :)<br>For the hydrofracturing calculations we refer to `Jourdain et al. 2025 <https://doi.org/10.5194/tc-19-1641-2025>`_ and the associated scripts: https://doi.org/10.5281/zenodo.13756240 and https://doi.org/10.5281/zenodo.15003864.</p>
Dataset and LOTOS program code to reproduce the main results of seismic tomography for West Aegean region (Turkey)
<p>This file contains the data to reproduce the results presented in the article: </p><p>Petrov I., Bushenkova N., Gulten, P., (2023). Intracontinental extension settings in the structure of the Aegean region (Turkey): local seismic tomography study., <i>Journal of Geodynamics.</i></p><p>This file includes:</p><p>1. The full version of the LOTOS code for the seismic tomography (Koulakov, 2009);</p><p>2. Folder with the dataset including arrival times of the P and S waves and the location of network from local seismicity in Aegean region of Tukey and it`s surroundings;</p><p>3. README.PDF file with the description of how to reproduce the tomography models and datatests based on data presented in the article. </p><p>Koulakov, I., 2009, LOTOS code for local earthquake tomographic inversion: Benchmarks for testing tomographic algorithms: Bulletin of the Seismological Society of America, v. 99, p. 194–214, https://doi.org/10.1785/0120080013.</p>
NLBayes Reproducibility Data Files
Open the record for dataset details and reuse information.
CBM algorithm TISMIR 2023: code and data for reproducing experiments
<p>This dataset contains the data necessary to reproduce experiments presented in the TISMIR article untitled "Barwise Music Structure Analysis with the Correlation Block-Matching Segmentation Algorithm" (under publication at the time of the upload, the link will be added after publication).</p><p>In details, this zenodo upload contains:</p><ul><li>Data, i.e. precomputed data and features (the self-similarity matrices in particular, along with beats and bars estimates) required to compute the CBM algorithm,</li><li>Code, i.e. the source files and the experiments (under the form of "Notebooks") used to compute results.</li></ul><p>This upload extends the version on git (https://gitlab.imt-atlantique.fr/a23marmo/autosimilarity_segmentation/-/tree/TISMIR).</p><p> </p>
Reproducibility of "Diffusion-based Generative AI for Exploring Transition States from 2D Molecular Graphs"
<p>This file is the source data to ensure reproducibility of the paper "Diffusion-based Generative AI for Exploring Transition States from 2D Molecular Graphs". It contains the logs and results of all DFT calculations associated with transition states generated using the model proposed in the paper. It also includes code to reproduce the core findings of the paper, which can be done by running reproduce.sh. To accurately reproduce the results of the paper, use the v1.0.0 virtual environment from "https://github.com/seonghann/tsdiff".</p>
EZBugs4Py: A benchmark of simple, easily reproducible Python bugs
<p>This is the appendix for paper entitled "EZBugs4Py: A benchmark of simple, easily reproducible Python bugs" submitted to MSR 2024.</p><p> </p><p>The dataset itself is available on <a href="https://github.com/gaborantal/ezbugs4py">GitHub.</a></p><p>This appendix contains:</p><ul><li>The results of GPT in the following tasks: automated program repair for to so-called "buggy" versions of the programs, the "failing" versions of the program, and the code synthesis based on the descriptions of the tasks.</li><li>The exact prompts we used in the paper.</li><li>The runner scripts to query GPT-4.</li><li>The categorization of the bugs.</li></ul><p> </p>
Code to reproduce Figures in "Adapting to Disruptions: Managing Supply Chain Resilience through Product Rerouting"
<p>This repository contains the code and data to reproduce the results of the paper:<em> "Adapting to Disruptions: Managing Supply Chain Resilience through Product Rerouting" </em>in<em> Science Advances</em>, https://doi.org/10.1126/sciadv.adj1194</p> <ul> <li>The `code/` directory contains the scripts used for the analysis and generating figures in the paper:</li> <li>The `data/` directory contains the data to create the plots</li> </ul> <p>The original data was obtained from <a href="https://www.slcg.com/opioid-data" target="_new">https://www.slcg.com/opioid-data</a>, which publicly distributes the Automation of Reports and Consolidated Orders System (ARCOS) dataset.</p>
Paper is not enough: Crowdsourcing the T<sub>1</sub> mapping common ground via the ISMRM reproducibility challenge
Dataset provided for NeuroLibre preprint. Author repo: https://www.github.com/rrsg2020/note NeuroLibre fork:https://github.com/roboneurolibre/note <p>For details, please visit the corresponding <a href="https://github.com/neurolibre/neurolibre-reviews/issues/23">NeuroLibre technical screening.</a></p> <p><strong><a href="https://neurolibre.org" target="NeuroLibre">https://neurolibre.org</a></strong></p>
Large-Scale Benchmarking of Metaphor-Based Optimization Heuristics -- Reproducibility Files
<p>This repository contains the files needed to reproduce the data, analysis and visualizations from the paper 'Large-Scale Benchmarking of Metaphor-Based Optimization Heuristics'.</p> <h3>Data collection</h3> <p>For the data collection, we make use of the IOHexperimenter library to access the BBOB functions. In addion, we rely on the set of libraries discussed in the paper. For the full set of libraries used during the benchmarking experiment, please refer to reqs_benchmark.txt (in the Benchmarking folder).</p> <p>The scripts used throughout the data collection are located in the 'benchmarking' directory. We have 1 file for each set of algorithms (expect mealpy and NIAPY, which are combined in the 'bestiary' file). Running these scripts generates a large amount of IOHprofiler-compatible files, which are availalbe in the 'Raw.zip' (Note: some of the libraries have their raw data split between multiple folders, if you want to run the full analysis, please merge these folders into one before processing). When running these scripts, please make sure to set the relevant FOLDER-variables at the top of the script!</p> <h3>Data processing</h3> <p>The raw data can be processed both for anytime performance (AOCC) and fixed-budget performance using the respective 'processing_METHOD.py' scripts. This generates several large csv-files, which are available in 'processed_data.zip'</p> <h3>Data analyis and visualization</h3> <p>The full set of analysis and corresponding visualizations can be reproduced using the 'Vizualization.ipynb' notebook. This notebook can use the processed performance data from the previous step, or the next level of processing with resulted in the remaining csv-files in this repository (this step is explained in the notebook itself).</p>
dominoSignal: data for a reproducible example
<p>This repository hosts example data for reproducible analysis of intra- and intercellular signaling in single cell RNA sequencing (scRNAseq) data based on transcription factor (TF) activation. We demonstrate analysis using dominoSignal on the 10X Genomics Peripheral Blood Mononuclear Cells (PBMC) data set of 2,700 cells <a href="https://cf.10xgenomics.com/samples/cell/pbmc3k/pbmc3k_filtered_gene_bc_matrices.tar.gz">PBMC3K</a>. scRNA-seq data is preprocessed following the <a href="https://satijalab.org/seurat/articles/pbmc3k_tutorial">Satija Lab's Guided Clustering Tutorial</a>. Quantification of TF activation is conducted using <a href="https://pyscenic.readthedocs.io/en/latest/">pySCENIC</a>. For more details on how this analysis is conducted, please refer to the vignettes in the <a href="https://github.com/FertigLab/dominoSignal" target="_blank" rel="noopener">dominoSignal package</a>.</p>
msiFlow: Automated Workflows for Reproducible and Scalable Multimodal Mass Spectrometry Imaging and Immunofluorescence Microscopy Data Processing and Analysis
<p>This record contains example and result data of msiFlow.</p> <p>msiFlow is a collection of automated workflows for reproducible and scalable multimodal mass spectrometry imaging (MSI) and immunofluorescence microscopy (IFM) data processing and analysis. Using an experimental mouse model for urinary tract infection, induced by uropathogenic E.coli (UPEC), we generated data by</p> <ul> <li>matrix-assisted laser desorption ionisation mass spectrometry imaging with laser-induced postionisation (MALDI-2 MSI) using the Bruker timsTOFfleX instrument</li> <li>transmission-mode MALDI-2 MSI (t-MALDI-2)</li> <li>immunofluorescence microscopy (IFM) using the MACSima system from Miltenyi </li> </ul> <p>msiFlow was tested on MALDI-2 MSI, t-MALDI-2 MSI and IFM data of control and UPEC-infected mouse bladder sections. In IFM we used Ly6G and actin for staining neutrophils and the muscle layer. We validated msiFlow on MALDI MSI data of bone marrow (BM)-derived neutrophils. Tentative lipid annotations were validated by MALDI DDA MSI and MALDI MS/MS. All data used and results generated by msiFlow are included in this dataset (besides the intermediate results of the MALDI-2 preprocessing due to data size).</p> <p>The dataset contains the following zip files:</p> <table> <tbody> <tr> <td><strong>zip file</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>ly6g_heterogeneity.zip</td> <td>example and result data (Ly6G clusters) for molecular_heterogeneity_flow</td> </tr> <tr> <td>if_segmentation.zip</td> <td>example and result data (Ly6G segmentation) for if_segmentation_flow</td> </tr> <tr> <td>ly6g_heterogeneity_signatures.zip</td> <td>example and result data (lipids for Ly6G clusters) for molecular_signatures_flow</td> </tr> <tr> <td>ly6g_molecular_signatures.zip</td> <td>example and result data (lipids for Ly6G) for molecular_signatures_flow</td> </tr> <tr> <td>msi_if_registration.zip</td> <td>example and result data for msi_if_registration_flow</td> </tr> <tr> <td>msi_segmentation.zip</td> <td>example and result data (segmented MSI bladder data) for msi_segmentation_flow</td> </tr> <tr> <td>region_group_analysis.zip</td> <td>example and result data (regulated lipids in different bladder tissue regions) for region_group_analysis_flow</td> </tr> <tr> <td>macsima.zip</td> <td>raw IFM data of UPEC-infected bladders containing Ly6G, actin and autofluorescence images</td> </tr> <tr> <td>maldi-bm-neutrophils.zip</td> <td>raw and pre-processed MALDI MSI data of BM-derived neutrophils</td> </tr> <tr> <td>t-maldi-2.zip</td> <td>raw t-MALDI-2 MSI data of a UPEC-infected bladder section</td> </tr> <tr> <td>maldi-2-<em>group-sampleno</em>.zip</td> <td>raw MALDI-2 MSI data of a control/UPEC bladder section</td> </tr> <tr> <td>MALDI_DDA_MSI.zip</td> <td>raw MALDI MSI data acquired in DDA mode</td> </tr> <tr> <td>TIMS_MS_MS.zip</td> <td>raw MALDI TIMS MS/MS data</td> </tr> </tbody> </table> <p> </p>
Data and code to reproduce analyses in Heinken et al, "A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites"
<p>This datasets archives the GitHub version found at https://github.com/ThieleLab/CodeBase to reproduce simulations for the article Heinken et al, "A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites", Cell Systems, in press.</p>
Reproducible analysis from SiRCle (Signature Regulatory Clustering)
<p>This contains the code and the data including the updates made during revisions for the manuscript: <strong><a href="https://www.biorxiv.org/content/10.1101/2022.07.02.498058v1.abstract">SiRCle (Signature Regulatory Clustering) model integration reveals mechanisms of phenotype regulation in renal cancer.</a> </strong></p> <p> </p> <p><strong>The data have been generated from CPTAC and TCGA.</strong> This includes no new data in this study.</p>
Reproducible Research in Archaeology
<p>This is a recording of a presentation given at the workshop on 'Reproducible research in Archaeology' at Durham University on 15th October 2021. The workshop included an introduction to what reproducibility is, why it is important for archaeological research and how you can make your research workflow reproducible. It also included some case studies demonstrating reproducible workflows used in archaeological research. This workshop was organised by Software Sustainability Institute Fellows - Alison Clarke and Emma Karoune. The slides from the presentation are available at <a href="https://doi.org/10.5281/zenodo.5564648">https://doi.org/10.5281/zenodo.5564648</a>.</p>
Alchemical Free Energy Estimators and Molecular Dynamics Engines: Accuracy, Precision and Reproducibility
<p>This zip contains all input structures for paper the: Alchemical Free<br> Energy Estimators and Molecular Dynamics<br> Engines: Accuracy, Precision and Reproducibility</p> <p>Authors: Alexander D. Wade, Agastya P. Bhati, Shunzhou Wan, Peter V.Coveney</p> <p>The structures of the folders are protein/ligand_transformation/alchemical_leg/input/files</p> <p>The ligand transformation are derived from previous work by wang et al. (https://pubs.acs.org/doi/10.1021/ja512751q)</p> <p>There are two files for the solvent alchemical leg: complex.pdb and complex.prmtop</p> <p>complex.pdb is structure file that also denotes the alchemical atoms in the pdb beta column. complex.prmtop is an AMBER parameter/topology file</p> <p>For the complex alchemical leg there is an additional file constraints.pdb that contains the constraint information in the pdb beta column.</p> <p>These files can be used with TIES_MD (https://ucl-ccs.github.io/TIES_MD/) or other molecular dynamics engiens that take AMBER input.</p>
Code and dataset used for the paper "The Economics of Research Reproducibility"
<p>This deposit includes the numerical data (code and dataset) used to generate the results displayed in the article "<a href="https://zenodo.org/record/5211312#.YRuv_I4zaUk">The Economics of Research Reproducibility</a>".</p>
Tissue-barriers-on-chip: paving the way for a reproducible practice in drug testing
<p>Supplementary material to the article</p> <p>Tissue-barriers-on-chip: paving the way for a reproducible practice in drug testing.</p> <p>Animation of the transfer- concept</p> <p>Cultivation on chip after transfer</p> <p>Cultivation on chip after transfer without air bubble measures</p> <p> </p> <p> </p>
SC20 Reproducibility Challenge Artifacts
<p>As a part of the reproducibility initiative at the SC20 conference, students participating in the Student Cluster Competition where given the reproducibility challenge, which is a directed reproducibility exercise that reproduces a paper from the previous year’s conference. At SC20, the paper, <em>MemXCT: memory-centric X-ray CT reconstruction with massive parallelization (</em><a href="https://doi.org/10.1145/3295500.3356220">https://doi.org/10.1145/3295500.3356220)</a>, from SC19 was used. In the directed exercise, students were asked to reproduce performance characteristics of the research on a variety of architectures. This file contains the experimental setup and outputs of their work from the reproducibility challenge. It also includes the direction that students were given to the students for the exercise.</p> <p>The students then published their reports as polished critiques of the reproducibility of the original paper along with a new revision of the original paper, titled <em>MemXCT: Design, Optimization, Scaling, and Reproducibility of X-Ray Tomography</em> <em>Imaging</em> (<a href="https://doi.org/10.1109/TPDS.2021.3128032">https://doi.org/10.1109/TPDS.2021.3128032</a>), in a Special Section of the IEEE Transactions on Parallel and Distributed Systems (TPDS). The paper along with the revised papers are available at <a href="https://ieeexplore.ieee.org/xpl/tocresult.jsp?isnumber=9696253&punumber=71"> https://ieeexplore.ieee.org/xpl/tocresult.jsp?isnumber=9696253&punumber=71</a>.</p>
"Constructing Temporal Networks of OSS Programming Language Ecosystems" reproducibility package
<p>Reproducibility package for the paper submission "Constructing Temporal Networks of OSS Programming Language Ecosystems". Contains the dataset, extracted and external metrics, and all scripts used for the construction and the analysis done in the paper.</p> <p> </p> <p>Anonymised for the double-blind review process.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.