Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
477
datasets available to search
ShareScore release 0.9.0
Dataset results
477 results for “input data”
ChemDyg input data
<p>Seven observations and reanalysis datasets used in the diagnostics sets of ChemDyg v1.0.0 are pre-processed, and the reference path is default assigned corresponding to different DOE machines. We provide the original data in NetCDF or text format. Please check https://github.com/E3SM-Project/ChemDyg for more information about ChemDyg. </p>
WaterFutures/BoN2024: BWDF input data and results
Open the record for dataset details and reuse information.
HANZE v2.3 flood impact model input data
<p>This dataset provides input data needed to run HANZE v2.3 model. The ZIP files need to be downloaded and unpacked in the same directory, which has to be defined in "get_file.py" of the HANZE model (variable "repo_path" at the beginning of the file).</p>
MUFFIN : A suite of tools for the analysis of functional sequencing data - Example input data
<p>This repository contains the data required to run the example notebooks and to reproduce the figures from the paper : </p> <div> <div><strong>MUFFIN : A suite of tools for the analysis of functional sequencing data</strong></div> </div> <div><em>Pierre de Langen, Benoit Ballester</em></div> <div>bioRxiv 2023.12.11.570597; doi: <a href="https://doi.org/10.1101/2023.12.11.570597" target="_blank" rel="noopener">https://doi.org/10.1101/2023.12.11.570597</a></div> <div> </div> <div>Source code is located here :</div> <div><a href="https://github.com/pdelangen/Muffin" target="_blank" rel="noopener">https://github.com/pdelangen/Muffin</a></div> <div> </div> <ul> <li><strong>10k_pbmc_gene/ </strong>contains the data for 10k pbmc dataset in standard 10x sparse count table format.</li> <li><strong>genome_annot/</strong> contains gencode v38 and chromosomes for the human (used for gene set enrichment analyses)</li> <li><strong>GO_files/</strong> contains gene set information retrieved from the g:ProfileR website.</li> <li><strong>immune_chip/ </strong>contains the data required to re-run the ChIP-seq analyses, it will require to also launch the dl_data.smk to retrieve the data from ENCODE.</li> <li><strong>tcga_atac/</strong> contains the sample-genomic region ATAC tag count table, as well as the sample metadata and a gene set file of cancer hallmark genes.</li> <li><strong>scATAC/</strong> contains the cell barcode-genomic region ATAC tag count table, as well as the barcode metadata and the 10k pbmc dataset pre-analyzed in h5 AnnData format.</li> </ul> <div> </div>
Input data and modelling files for a model of the Finnish energy system with focus on cascade hydropower and the addition of a hydrogen storage system realised in Backbone
<p>The files show the input data and modelling files used for the publication "Cascade hydropower integration in a techno-economic power system model: A study of Finnish hydropower plants" (Kiehle et al., 2025 - submitted). The paper's <a title="Preprint on SSRN" href="https://dx.doi.org/10.2139/ssrn.4971685" target="_blank" rel="noopener">preprint</a> is available. A model of the Finnish energy system in 2022 was built in the techno-economic modelling framework Backbone (available on GitLab: https://gitlab.vtt.fi/backbone/backbone). The focus was on implementing cascading hydropower plants in a power system model, including individual reservoirs, generation and spillage capacities. </p> <p>"ModellingFiles_Debug" are GAMS-based data that can be used to run the scenario in Backbone or display the results. "ModellingResults" are gdx files that purely list the results. Those are also presented in more detail in the scientific paper. The Excel files present the input data used for modelling and can also be used to run the model. </p>
Example data sets and input parameters for running various features in the Python code Dynpy
<p>This is a set of examples intended to be used with the Python code called Dynpy at <a href="https://zenodo.org/records/13241475">https://zenodo.org/records/13241475</a></p>
Input and Output Data for local earthquake tomography in the central Dead Sea Fault using PyVoroTomo
<p>Data to reproduce wave velocity models for the central DSF.</p> <p>eventsAndArrivals.h5 - 2 csv files (keys: events, arrivals)</p> <p>stations_sub.h5 - csv file containing station data</p> <p>gitter.csv - 1D velocity model by Gitterman et al. (2002)</p> <p>3d_Vp_Vs_VpVs_models.nc - velocity models for vp, vs, and vp/vs and their uncertainty.</p> <p>relocated seismicity.h5 - relocated seismicity, arrivals used, and stations used (keys: events, arrivals, stations)</p>
Input data and code supporting the cod_v2 population estimates
<p>The <strong><em>model.zip</em></strong> file contains input data and code supporting the cod_v2 population estimates. The file<strong> <em>modelData.RData</em></strong> provides the input data to the JAGS model and the file<em> <strong>modelCode.R</strong></em> contains the source code for the model in the JAGS language. The files can be used to run the model for further assessments and as a starting point for further model development.</p> <p>The data and the model were developed using the statistical software <strong>R version 4.0.2</strong> (https://cran.r-project.org/bin/windows/base/old/4.0.2) and <strong>JAGS 4.3.0</strong> (https://mcmc-jags.sourceforge.io), a program for analysis of Bayesian graphical models using Gibbs sampling, through the R package <strong>runjags 2.2.0</strong> (https://cran.r-project.org/web/packages/runjags).</p>
Crossbow input data for ATC and FB aproach
<p>These data refears to the load, generation and already allocated capacities (AAC) for five countries in the South Eastern Europe. Also these data are reported for specific days in hourly basis as they are retrieved from Entso-E tranasparency platform. Simulation scenarios took place using this data for the ATC and FB aproach in SEE region in the frame of Crossbow project.</p>
Combined data file for Jilbert et al. "Anthropogenic Inputs of Terrestrial Organic Matter Influence Carbon Loading and Methanogenesis in Coastal Baltic Sea Sediments", Frontiers in Earth Science 9, 2021
<p>The datafile contains all the new raw data presented in the figures in the publication.</p>
Carbon sequestration of a forested wetland receiving nutrient inputs - soil, tree and greenhouse gas data
<p><span><span><span><span><span><span><span><span><span><span><span>Here we describe a pilot wetland carbon project located 30 km west of New Orleans where measurements were taken in 2013 and 2018, and applied to the carbon offset methodology, "Restoration of Degraded Deltaic Wetlands of the Mississippi Delta" ("the ACR Methodology") published by the American Carbon Registry (ACR). Baseline emissions were modeled using values derived from scientific literature. Results indicate net sequestration rate of 619,727 tons carbon dioxide equivalent (CO<sub>2</sub>e) over the 40 year project duration, which equates to 16,527 t CO2-e/yr, if wetland greenhouse gases (GHGs) are included, and 200,143 t CO<sub>2</sub>e over 40 years, or 5,003 t CO2-e/yr, if wetland greenhouse gasses were conservatively omitted. A kriging exercise was carried out that modeled the tree and soil pools, which resulted in net sequestration of 723,375 t CO2-e over 40 years (annual mean 18,084 t CO2-e/yr) with greenhouse gases, and 262,472 t CO2-e over 40 years (annual mean rate 6,560 t CO2-e/yr) if greenhouse gases were omitted. Unfortunately, the project was withdrawn, prohibiting the issuance and eventual transaction of carbon credits, due to very large uncertainty estimates mostly associated with GHG emissions and the kriging approach as in situ sampling could not be conducted as required by the methodology.</span></span></span></span></span></span></span></span></span></span></span></p>
The ECS-FVCOM results and also the input variables of FMGDM, the initial particles, satellite pictures, and trajectories data
<p>The data for the manuscript: A Lagrangian-based Floating Macro<br> -algal Growth and Drift Model (FMGDM v1.0): application to the Yellow Sea green tides. <br> -----------------------------------------------------------------------------<br> Updata: 2021/03/18</p> <p>-----------------------------------------------------------------------------<br> Part 1: The ECS-FVCOM results and also the input variables of FMGDM<br> Part 2: The initial particles position imformation (lat, lon, depth)<br> Part 3: The satellite pictures of green tides in YS, 2014 and 2015</p> <p>Part 4: Drifters trajectories dataset (lat, lon)</p> <p> </p> <p>=============================================</p> <p>Version 5 updata: 2021/12/20</p> <p>Note: Modified and added some missing variables ('omega') in Part1,</p> <p> the results of ECS-FVCOM</p>
Simulation Input Data for "Molecular simulation of lignin-related aromatic compound permeation through Gram-negative bacterial outer membranes"
<p>This is the reduced data behind an upcoming manuscript investigating permeability across the outer membranes of Gram-negative bacteria. The data is taken directly from the directory structure that contains both the simulation and analysis, with excluded trajectory files and intermediate products to fit within the zenodo upload limit. The tar command used to generate this tarball was:</p> <pre><code class="language-bash">tar --exclude="*BAK" --exclude="*dcd" --exclude="*xsc" --exclude="*vel" --exclude="*coor" --exclude="*csv" --exclude="*old" --exclude="*log" --exclude="*state" --exclude="*watpos/*npz" --exclude="*new*png" --exclude="*frame*png" --exclude="*ppm" --exclude="*mp4" --exclude="*bayesdata*npy" --exclude="*run.npy" -zcvf OM.tgz OuterMembrane</code></pre> <p>Within the OuterMembrane directory, there are 3 primary subdirectories.</p> <ul> <li><strong>Build </strong>contains the scripts and files to build the simulation systems, including the CHARMM-GUI output</li> <li><strong>Equilibrium</strong> contains the equilibrium simulation inputs and the analysis scripts (subdirectory <strong>Analysis</strong>)</li> <li><strong>REUS2</strong>, which has the replica exchange inputs and essential output. It also contains an <strong>Analysis</strong> subdirectory that carries out the analysis within the text.</li> </ul>
Input data for Narayan et al 2022.
<p>Please use this input data to reproduce the results, figures for <strong>Evaluation of uncertainties in the anthropogenic SO2 emissions in the USA from NASA’s OMI point source catalog - </strong>by<strong> </strong>Kanishka Narayan, Steven J. Smith, Vitali E. Fioletov<sup> </sup>& Chris A. McLinden<span>, 2022</span></p>
Input data for Differential NicheNet analysis performed in the liver atlas paper Guilliams et al., Cell 2022
<p>Input data for Differential NicheNet analysis performed in the liver atlas paper Guilliams et al., Cell 2022</p> <p>See https://github.com/saeyslab/NicheNet_LiverCellAtlas and https://www.sciencedirect.com/science/article/pii/S0092867421014811</p>
Data from: Effects of input data sources on species distribution model predictions across species with different distributional ranges
<p>Species distribution models (SDMs) are a popular tool in theoretical and quantitative ecology, and constitute the most widely used modelling framework in global change science and biodiversity conservation. As main data sources, SDMs require georeferenced biodiversity observations as a response or dependent variable (e.g. species occurrence, species richness, etc) and geographic layers of environmental information as predictors or independent variables (e.g. climate, land cover, vegetation indices derived from remote sensing, etc). However, although SDMs have become one of the most important quantitative tools for addressing regular and timely biodiversity assessments worldwide, these techniques are still subject to different sources of uncertainty that have been unequally assessed. Thus, despite uncertainty related to niche-based or distribution-based models has been addressed at different stages in the modelling process, an analysis of the effect of uncertainty coming from alternative data sources on the predictive ability of SDMs is still limited.</p> <p>Citizen-collected species occurrence data (e.g. eBird) are often used for fitting SDMs when data from standardized and expert-supported surveys (e.g. Atlases) are unavailable. On the other hand, macroclimate variables are much more commonly used as predictors in SDMs than other sources of information coming from remote sensing data. We assessed the effects of using different data sources (in both response and predictor variables) on SDM performance across a wide range of bird species with contrasting distributional ranges in the Iberian Peninsula (Portugal and Spain). To do that, a SDM ensemble-forecasting approach was implemented by using bird data from two different data sources: the semi-structured eBird project and standardized Atlases. We fitted SDMs with three predictor types: macroclimate, remotely sensed ecosystem functional attributes (EFAs) from vegetation indices, and their combination. Species were grouped in four range size classes. We also used different evaluation metrics to better assess the uncertainty of model predictions. We then applied generalized linear mixed-effects models to test the effect on model performance of input data source across distributional range sizes while accounting for different accuracy metrics. Pairwise comparisons between range projections were used to assess their spatial similarity.</p> <p>Our models demonstrated the usefulness and complementarity of different input data sources when modelling species distribution across different distributional ranges. Citizen science and remote sensing data contribute to update the knowledge of the distribution of the most threatened bird species by increasing the model accuracy. These findings highlight the need to integrate different data sources to improve the model predictions at regional scale. Our framework also underlines that model uncertainty should be examined more exhaustively at early stages of the modelling process.</p> <p>To perfom and replicate this study, this dataset provides all needed files (as tables) to fit SDMs: i) the Iberian bird species occurrences at 10km UTM square as a response or dependent variable; ii) the geographic layers of environmental information at 10km UTM square for the Iberian Peninsula as predictors or independent variables, such as climate data, ecosystem functioning attributes (EFAs) and the combined climate and EFA data. The dataset is provided by four <em>*.csv</em> files named as:</p> <p><em>1) The_Iberian_bird_species_occurrences_dataset_10km.csv</em></p> <p><em>2) CHELSA_bioclimate_variables_IP10km.csv</em></p> <p><em>3) MODIS_EVI-based_EFAs_IP10km.csv</em></p> <p><em>4) Combined_bioclimate_EFA_dataset_IP10km.csv</em></p> <p>For a more detailed description of the main dataset and each of these subdatasets, please refer to the attached README file.</p> <p><strong>Keywords:</strong> bird atlas, eBird data, ecosystem functional attributes (EFAs), Iberian Peninsula, IUCN categories, Model accuracy, MODIS EVI, narrow-ranged species, remote sensing, species distribution models (SDMs), widespread species</p>
Input features and benchmark data sets for protein complex prediction and E. coli proteome application by AF2Complex
<p>Benchmark data sets of AF2Complex, input features for application to E. coli proteome, and predicted structural models of E. coli Ccm I as described in</p> <p><strong>Predicting direct physical interactions in multimeric proteins with deep learning</strong></p> <p><em>Mu Gao, Davi Nakajima An, Jerry M. Parks, Jeffrey Skolnick</em></p> <ol> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/af2complex_bench.tar.gz">af2complex_bench.tar.gz</a>: Benchmark data sets CP17, Dimer1193 and Oligomer562, including input features for AF2Complex/AF-Multimer, both paired and unpaired MSAs, as well as sequences, experimental structures, and results presented in the AF2Complex work (~90GB de-compressed size)</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_Ccm_I.tar.gz">ecoli_Ccm_I.tar.gz</a>: Computational models of the<em> E. coli</em> Ccm I system</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_set.tar.gz">ecoli_sets.tar.gz</a>: Lists of benchmark sets of positive and negative PPIs from <em>E. coli</em></li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_af_fea.tar.gz">ecoli_af_fea.tar.gz: </a>Pre-generated input features of <em>E. coli</em> proteome for protein complex prediction and modeling by AF2Complex (4,429 proteins, ~800 GB de-compressed size). This data set can be used with AF2Complex to probe the interactions of any combinations among the 4,429 proteins of E. coli.</li> </ol> <p> </p>
AirGAM 2022r1 input data for all stations 2005-2019
<p>Contains EEA Airbase/AQ e-Reporting data and ECMWF ERA 5 met data for all stations.</p>
Data from: Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range - wild grapevine sampling locations, Maxent input files, morphological and microsatellite data
<p><span>This dataset contains raw data described in the paper: "Rahimi O., Ohana-Levi N., Brauner H., Inbar N., Hübner S. and Drori E. (2021) "Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range", accepted for publication in "Ecology and Evolution".</span></p> <p><span>The spatial distribution of plants is constrained by demographic and eco-geographic factors that determine the range and abundance of the species. In this study, we performed genetic and morphological analyzes based on SSR and OIV datasets. In addition, according to the spatial distribution model performed by Maxent software we found that distance to water sources, Normalized difference vegetation index, and precipitation are the main environmental factors constraining <i>V.v. sylvestris</i> distribution at its southern distribution range. All raw data used for this study can be found in this deposit which contains a table with grapevine locations, Maxent input files, morphological and microsatellite data. </span></p>
Input files and data for sensitivity analysis of ParFlow-CLM
<p>This zip file includes all the input files and forcing data for the sensitivity analysis of ParFlow-CLM.</p> <p>The file of "Stettbach_SA" includes all the information for the sensitivity analysis of Stettbach catchment.</p> <p>The file of "high_slope_SA" includes all the information for the sensitivity analysis of high slope effect in section 2.5.1 Topography effects.</p> <p>The file of "low_slope_SA" includes all the information for the sensitivity analysis of low slope effect in section 2.5.1 Topography effects.</p> <p>The file of "Barcelona_SA" includes all the information for the sensitivity analysis of Barcelona in section 2.5.2 Climate effects.</p> <p>The file of "LK_SA" includes all the information for the sensitivity analysis of Lemmenjoke national park in Finland in section 2.5.2 Climate effects.</p> <p>The file of "Stettbach_2020_2021" includes all the information for the runoff comparison between simulation and observation in 2020 and 2021 in section 3.1 (Model performance in Stettbach catchment).</p> <p>The file of "results_process_code" includes all the scripts used for the analysis of this study.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.