Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

477

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

477 results for “input data”

Learn how ShareScore rates datasets ↗
zenodo36/100

ChemDyg input data

<p>Seven observations and reanalysis datasets used in the diagnostics sets of ChemDyg v1.0.0 are pre-processed, and the reference path is default assigned corresponding to different DOE machines. We provide the original data in NetCDF or text format. Please check https://github.com/E3SM-Project/ChemDyg for more information about ChemDyg.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

WaterFutures/BoN2024: BWDF input data and results

Open the record for dataset details and reuse information.

openapache2.0Mar 2024View details →
zenodo36/100

HANZE v2.3 flood impact model input data

<p>This dataset provides input data needed to run HANZE v2.3 model. The ZIP files need to be downloaded and unpacked in the same directory, which has to be defined in "get_file.py" of the HANZE model (variable "repo_path" at the beginning of the file).</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

MUFFIN : A suite of tools for the analysis of functional sequencing data - Example input data

<p>This repository contains the data required to run the example notebooks and to reproduce the figures from the paper :&nbsp;</p> <div> <div><strong>MUFFIN : A suite of tools for the analysis of functional sequencing data</strong></div> </div> <div><em>Pierre&nbsp;de Langen,&nbsp;Benoit&nbsp;Ballester</em></div> <div>bioRxiv&nbsp;2023.12.11.570597;&nbsp;doi:&nbsp;<a href="https://doi.org/10.1101/2023.12.11.570597" target="_blank" rel="noopener">https://doi.org/10.1101/2023.12.11.570597</a></div> <div>&nbsp;</div> <div>Source code is located here :</div> <div><a href="https://github.com/pdelangen/Muffin" target="_blank" rel="noopener">https://github.com/pdelangen/Muffin</a></div> <div>&nbsp;</div> <ul> <li><strong>10k_pbmc_gene/ </strong>contains the data for 10k pbmc dataset in standard 10x sparse count table format.</li> <li><strong>genome_annot/</strong> contains gencode v38 and chromosomes for the human (used for gene set enrichment analyses)</li> <li><strong>GO_files/</strong> contains gene set information retrieved from the g:ProfileR website.</li> <li><strong>immune_chip/ </strong>contains the data required to re-run the ChIP-seq analyses, it will require to also launch the dl_data.smk to retrieve the data from ENCODE.</li> <li><strong>tcga_atac/</strong> contains the sample-genomic region ATAC tag count table, as well as the sample metadata and a gene set file of cancer hallmark genes.</li> <li><strong>scATAC/</strong> &nbsp;contains the cell barcode-genomic region ATAC tag count table, as well as the barcode metadata and the 10k pbmc dataset pre-analyzed in h5 AnnData format.</li> </ul> <div>&nbsp;</div>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Input data and modelling files for a model of the Finnish energy system with focus on cascade hydropower and the addition of a hydrogen storage system realised in Backbone

<p>The files show the input data and modelling files used for the publication "Cascade hydropower integration in a techno-economic power system model: A study of Finnish hydropower plants" (Kiehle et al., 2025 - submitted). The paper's <a title="Preprint on SSRN" href="https://dx.doi.org/10.2139/ssrn.4971685" target="_blank" rel="noopener">preprint</a> is available. A model of the Finnish energy system in 2022 was built in the techno-economic modelling framework Backbone (available on GitLab: https://gitlab.vtt.fi/backbone/backbone). The focus was on implementing cascading hydropower plants in a power system model, including individual reservoirs, generation and spillage capacities.&nbsp;</p> <p>"ModellingFiles_Debug" are GAMS-based data that can be used to run the scenario in Backbone or display the results. "ModellingResults" are gdx files that purely list the results. Those are also presented in more detail in the scientific paper. The Excel files present the input data used for modelling and can also be used to run the model.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Example data sets and input parameters for running various features in the Python code Dynpy

<p>This is a set of examples intended to be used with the Python code called Dynpy at <a href="https://zenodo.org/records/13241475">https://zenodo.org/records/13241475</a></p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Input and Output Data for local earthquake tomography in the central Dead Sea Fault using PyVoroTomo

<p>Data to reproduce wave velocity models for the central DSF.</p> <p>eventsAndArrivals.h5 - 2 csv files (keys: events, arrivals)</p> <p>stations_sub.h5 - csv file containing station data</p> <p>gitter.csv - 1D velocity model by Gitterman et al. (2002)</p> <p>3d_Vp_Vs_VpVs_models.nc - velocity models for vp, vs, and vp/vs and their uncertainty.</p> <p>relocated seismicity.h5 - relocated seismicity, arrivals used, and stations used (keys: events, arrivals, stations)</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Input data and code supporting the cod_v2 population estimates

<p>The <strong><em>model.zip</em></strong> file contains input&nbsp;data and code supporting the&nbsp;cod_v2 population estimates.&nbsp;The file<strong> <em>modelData.RData</em></strong> provides the input data to the JAGS model and the file<em> <strong>modelCode.R</strong></em> contains the source code for the model in the JAGS language. The files can be used to run the model for further assessments and as a starting point for further model development.</p> <p>The data and the model were developed using the statistical software <strong>R version 4.0.2</strong>&nbsp;(https://cran.r-project.org/bin/windows/base/old/4.0.2) and <strong>JAGS 4.3.0</strong> (https://mcmc-jags.sourceforge.io), a program for analysis of Bayesian graphical models using Gibbs sampling, through the R package <strong>runjags 2.2.0</strong> (https://cran.r-project.org/web/packages/runjags).</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Crossbow input data for ATC and FB aproach

<p>These&nbsp;data refears to the load, generation and already allocated capacities (AAC) for five countries in the South&nbsp;Eastern Europe. Also these&nbsp;data are reported for&nbsp;specific days&nbsp;in hourly basis as they are retrieved from Entso-E tranasparency platform. Simulation scenarios took place using this data for the ATC and FB aproach in SEE region in the frame of Crossbow project.</p>

opencc-by-4.0Mar 2021View details →
zenodo36/100

Combined data file for Jilbert et al. "Anthropogenic Inputs of Terrestrial Organic Matter Influence Carbon Loading and Methanogenesis in Coastal Baltic Sea Sediments", Frontiers in Earth Science 9, 2021

<p>The datafile contains all the new raw data presented in the figures in the publication.</p>

opencc-by-4.0Nov 2021View details →
dryad36/100

Carbon sequestration of a forested wetland receiving nutrient inputs - soil, tree and greenhouse gas data

<p><span><span><span><span><span><span><span><span><span><span><span>Here we describe a pilot wetland carbon project located 30 km west of New Orleans where measurements were taken in 2013 and 2018, and applied to the carbon offset methodology, "Restoration of Degraded Deltaic Wetlands of the Mississippi Delta" ("the ACR Methodology") published by the American Carbon Registry (ACR). Baseline emissions were modeled using values derived from scientific literature. Results indicate net sequestration rate of 619,727 tons carbon dioxide equivalent (CO<sub>2</sub>e) over the 40 year project duration, which equates to 16,527 t CO2-e/yr, if wetland greenhouse gases (GHGs) are included, and 200,143 t CO<sub>2</sub>e over 40 years, or 5,003 t CO2-e/yr, if wetland greenhouse gasses were conservatively omitted. A kriging exercise was carried out that modeled the tree and soil pools, which resulted in net sequestration of 723,375 t CO2-e over 40 years (annual mean 18,084 t CO2-e/yr) with greenhouse gases, and 262,472 t CO2-e over 40 years (annual mean rate 6,560 t CO2-e/yr) if greenhouse gases were omitted. Unfortunately, the project was withdrawn, prohibiting the issuance and eventual transaction of carbon credits, due to very large uncertainty estimates mostly associated with GHG emissions and the kriging approach as in situ sampling could not be conducted as required by the methodology.</span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroDec 2021View details →
zenodo36/100

The ECS-FVCOM results and also the input variables of FMGDM, the initial particles, satellite pictures, and trajectories data

<p>The data for&nbsp;the manuscript:&nbsp;A Lagrangian-based Floating Macro<br> -algal Growth and Drift Model (FMGDM v1.0): application to the&nbsp;Yellow Sea green tides.&nbsp;<br> -----------------------------------------------------------------------------<br> Updata: 2021/03/18</p> <p>-----------------------------------------------------------------------------<br> Part 1: The ECS-FVCOM results and also the input variables of FMGDM<br> Part 2: The initial particles position imformation (lat, lon, depth)<br> Part 3: The satellite pictures of green tides in YS, 2014 and 2015</p> <p>Part 4: Drifters trajectories dataset (lat, lon)</p> <p>&nbsp;</p> <p>=============================================</p> <p>Version 5 updata: 2021/12/20</p> <p>Note: Modified and added some missing variables (&#39;omega&#39;) in Part1,</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; the results of ECS-FVCOM</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Simulation Input Data for "Molecular simulation of lignin-related aromatic compound permeation through Gram-negative bacterial outer membranes"

<p>This is the reduced data behind an upcoming manuscript investigating permeability across the outer membranes of Gram-negative bacteria. The data is taken directly from the directory structure that contains both the simulation and analysis, with excluded trajectory files and intermediate products to fit within the zenodo upload limit. The tar command used to generate this tarball was:</p> <pre><code class="language-bash">tar --exclude="*BAK" --exclude="*dcd" --exclude="*xsc" --exclude="*vel" --exclude="*coor" --exclude="*csv" --exclude="*old" --exclude="*log" --exclude="*state" --exclude="*watpos/*npz" --exclude="*new*png" --exclude="*frame*png" --exclude="*ppm" --exclude="*mp4" --exclude="*bayesdata*npy" --exclude="*run.npy" -zcvf OM.tgz OuterMembrane</code></pre> <p>Within the OuterMembrane directory, there are 3 primary subdirectories.</p> <ul> <li><strong>Build </strong>contains the scripts and files to build the simulation systems, including the CHARMM-GUI output</li> <li><strong>Equilibrium</strong> contains the equilibrium simulation inputs and the analysis scripts (subdirectory <strong>Analysis</strong>)</li> <li><strong>REUS2</strong>, which has the replica exchange inputs and essential output. It also contains an <strong>Analysis</strong> subdirectory that carries out the analysis within the text.</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Input data for Narayan et al 2022.

<p>Please use this input data to reproduce the results, figures for&nbsp;<strong>Evaluation of uncertainties in the anthropogenic SO2 emissions in the USA from NASA&rsquo;s OMI point source catalog - </strong>by<strong>&nbsp;</strong>Kanishka Narayan, Steven J. Smith, Vitali E. Fioletov<sup>&nbsp;</sup>&amp; Chris A. McLinden<span>, 2022</span></p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Input data for Differential NicheNet analysis performed in the liver atlas paper Guilliams et al., Cell 2022

<p>Input data for Differential NicheNet analysis performed in the liver atlas paper Guilliams et al., Cell 2022</p> <p>See&nbsp;https://github.com/saeyslab/NicheNet_LiverCellAtlas and&nbsp;https://www.sciencedirect.com/science/article/pii/S0092867421014811</p>

opencc-by-4.0Jan 2022View details →
dryad36/100

Data from: Effects of input data sources on species distribution model predictions across species with different distributional ranges

<p>Species distribution models (SDMs) are a popular tool in theoretical and quantitative ecology, and constitute the most widely used modelling framework in global change science and biodiversity conservation. As main data sources, SDMs require georeferenced biodiversity observations as a response or dependent variable (e.g. species occurrence, species richness, etc) and geographic layers of environmental information as predictors or independent variables (e.g. climate, land cover, vegetation indices derived from remote sensing, etc). However, although SDMs have become one of the most important quantitative tools for addressing regular and timely biodiversity assessments worldwide, these techniques are still subject to different sources of uncertainty that have been unequally assessed. Thus, despite uncertainty related to niche-based or distribution-based models has been addressed at different stages in the modelling process, an analysis of the effect of uncertainty coming from alternative data sources on the predictive ability of SDMs is still limited.</p> <p>Citizen-collected species occurrence data (e.g. eBird) are often used for fitting SDMs when data from standardized and expert-supported surveys (e.g. Atlases) are unavailable. On the other hand, macroclimate variables are much more commonly used as predictors in SDMs than other sources of information coming from remote sensing data. We assessed the effects of using different data sources (in both response and predictor variables) on SDM performance across a wide range of bird species with contrasting distributional ranges in the Iberian Peninsula (Portugal and Spain). To do that, a SDM ensemble-forecasting approach was implemented by using bird data from two different data sources: the semi-structured eBird project and standardized Atlases. We fitted SDMs with three predictor types: macroclimate, remotely sensed ecosystem functional attributes (EFAs) from vegetation indices, and their combination. Species were grouped in four range size classes. We also used different evaluation metrics to better assess the uncertainty of model predictions. We then applied generalized linear mixed-effects models to test the effect on model performance of input data source across distributional range sizes while accounting for different accuracy metrics. Pairwise comparisons between range projections were used to assess their spatial similarity.</p> <p>Our models demonstrated the usefulness and complementarity of different input data sources when modelling species distribution across different distributional ranges. Citizen science and remote sensing data contribute to update the knowledge of the distribution of the most threatened bird species by increasing the model accuracy. These findings highlight the need to integrate different data sources to improve the model predictions at regional scale. Our framework also underlines that model uncertainty should be examined more exhaustively at early stages of the modelling process.</p> <p>To perfom and replicate this study, this dataset provides all needed files (as tables) to fit SDMs: i) the Iberian bird species occurrences at 10km UTM square as a response or dependent variable;  ii) the geographic layers of environmental information at 10km UTM square for the Iberian Peninsula as predictors or independent variables, such as climate data, ecosystem functioning attributes (EFAs) and the combined climate and EFA data. The dataset is provided by four <em>*.csv</em> files named as:</p> <p><em>1) The_Iberian_bird_species_occurrences_dataset_10km.csv</em></p> <p><em>2) CHELSA_bioclimate_variables_IP10km.csv</em></p> <p><em>3) MODIS_EVI-based_EFAs_IP10km.csv</em></p> <p><em>4) Combined_bioclimate_EFA_dataset_IP10km.csv</em></p> <p>For a more detailed description of the main dataset and each of these subdatasets, please refer to the attached README file.</p> <p><strong>Keywords:</strong> bird atlas, eBird data, ecosystem functional attributes (EFAs), Iberian Peninsula, IUCN categories, Model accuracy, MODIS EVI, narrow-ranged species, remote sensing, species distribution models (SDMs), widespread species</p>

opencc-zeroFeb 2022View details →
zenodo36/100

Input features and benchmark data sets for protein complex prediction and E. coli proteome application by AF2Complex

<p>Benchmark data sets of AF2Complex, input features for application to E. coli proteome, and predicted structural models of E. coli Ccm I as described in</p> <p><strong>Predicting direct physical interactions in multimeric proteins with deep learning</strong></p> <p><em>Mu Gao, Davi Nakajima An, Jerry M. Parks, Jeffrey Skolnick</em></p> <ol> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/af2complex_bench.tar.gz">af2complex_bench.tar.gz</a>:&nbsp; Benchmark data sets CP17, Dimer1193 and Oligomer562, including input features for AF2Complex/AF-Multimer, both paired and unpaired MSAs, as well as sequences, experimental structures, and results presented in the AF2Complex work (~90GB de-compressed size)</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_Ccm_I.tar.gz">ecoli_Ccm_I.tar.gz</a>: Computational models of the<em> E. coli</em> Ccm I system</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_set.tar.gz">ecoli_sets.tar.gz</a>: Lists of benchmark sets of positive and negative PPIs from <em>E. coli</em></li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_af_fea.tar.gz">ecoli_af_fea.tar.gz: </a>Pre-generated input features of <em>E. coli</em> proteome for protein complex prediction and modeling by AF2Complex (4,429 proteins, ~800 GB de-compressed size). This data set can be used with AF2Complex to probe the interactions of any combinations among the 4,429 proteins of E. coli.</li> </ol> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

AirGAM 2022r1 input data for all stations 2005-2019

<p>Contains EEA Airbase/AQ e-Reporting data&nbsp;and ECMWF ERA 5 met data for all stations.</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

Data from: Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range - wild grapevine sampling locations, Maxent input files, morphological and microsatellite data

<p><span>This dataset contains raw data described in the paper: "Rahimi O., Ohana-Levi N., Brauner H., Inbar N., Hübner S. and Drori E. (2021) "Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range",  accepted for publication in "Ecology and Evolution".</span></p> <p><span>The spatial distribution of plants is constrained by demographic and eco-geographic factors that determine the range and abundance of the species. In this study, we performed genetic and morphological analyzes based on SSR and OIV datasets. In addition, according to the spatial distribution model performed by Maxent software we found that distance to water sources, Normalized difference vegetation index, and precipitation are the main environmental factors constraining <i>V.v. sylvestris</i> distribution at its southern distribution range. All raw data used for this study can be found in this deposit which contains a table with grapevine locations, Maxent input files, morphological and microsatellite data. </span></p>

opencc-zeroApr 2022View details →
zenodo36/100

Input files and data for sensitivity analysis of ParFlow-CLM

<p>This zip file includes all the input files and forcing data for the sensitivity analysis of ParFlow-CLM.</p> <p>The file of &quot;Stettbach_SA&quot; includes all the information for the&nbsp;sensitivity analysis of Stettbach catchment.</p> <p>The file of &quot;high_slope_SA&quot; includes&nbsp;all the information for the&nbsp;sensitivity analysis of high slope effect&nbsp;in section 2.5.1&nbsp;Topography effects.</p> <p>The file of &quot;low_slope_SA&quot; includes&nbsp;all the information for the&nbsp;sensitivity analysis of low slope effect&nbsp;in section 2.5.1&nbsp;Topography effects.</p> <p>The file of &quot;Barcelona_SA&quot; includes&nbsp;all the information for the&nbsp;sensitivity analysis of Barcelona in section 2.5.2 Climate effects.</p> <p>The file of &quot;LK_SA&quot; includes&nbsp;all the information for the&nbsp;sensitivity analysis of Lemmenjoke national park in&nbsp;Finland in section 2.5.2 Climate effects.</p> <p>The file of &quot;Stettbach_2020_2021&quot; includes&nbsp;all the information for the&nbsp;runoff comparison between simulation and observation in&nbsp;2020 and 2021 in section 3.1 (Model performance in Stettbach catchment).</p> <p>The file of &quot;results_process_code&quot; includes all the scripts used for the analysis of this study.</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record