Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,600
datasets available to search
ShareScore release 0.9.0
Dataset results
1,600 results for “input”
Data from: Effects of input data sources on species distribution model predictions across species with different distributional ranges
<p>Species distribution models (SDMs) are a popular tool in theoretical and quantitative ecology, and constitute the most widely used modelling framework in global change science and biodiversity conservation. As main data sources, SDMs require georeferenced biodiversity observations as a response or dependent variable (e.g. species occurrence, species richness, etc) and geographic layers of environmental information as predictors or independent variables (e.g. climate, land cover, vegetation indices derived from remote sensing, etc). However, although SDMs have become one of the most important quantitative tools for addressing regular and timely biodiversity assessments worldwide, these techniques are still subject to different sources of uncertainty that have been unequally assessed. Thus, despite uncertainty related to niche-based or distribution-based models has been addressed at different stages in the modelling process, an analysis of the effect of uncertainty coming from alternative data sources on the predictive ability of SDMs is still limited.</p> <p>Citizen-collected species occurrence data (e.g. eBird) are often used for fitting SDMs when data from standardized and expert-supported surveys (e.g. Atlases) are unavailable. On the other hand, macroclimate variables are much more commonly used as predictors in SDMs than other sources of information coming from remote sensing data. We assessed the effects of using different data sources (in both response and predictor variables) on SDM performance across a wide range of bird species with contrasting distributional ranges in the Iberian Peninsula (Portugal and Spain). To do that, a SDM ensemble-forecasting approach was implemented by using bird data from two different data sources: the semi-structured eBird project and standardized Atlases. We fitted SDMs with three predictor types: macroclimate, remotely sensed ecosystem functional attributes (EFAs) from vegetation indices, and their combination. Species were grouped in four range size classes. We also used different evaluation metrics to better assess the uncertainty of model predictions. We then applied generalized linear mixed-effects models to test the effect on model performance of input data source across distributional range sizes while accounting for different accuracy metrics. Pairwise comparisons between range projections were used to assess their spatial similarity.</p> <p>Our models demonstrated the usefulness and complementarity of different input data sources when modelling species distribution across different distributional ranges. Citizen science and remote sensing data contribute to update the knowledge of the distribution of the most threatened bird species by increasing the model accuracy. These findings highlight the need to integrate different data sources to improve the model predictions at regional scale. Our framework also underlines that model uncertainty should be examined more exhaustively at early stages of the modelling process.</p> <p>To perfom and replicate this study, this dataset provides all needed files (as tables) to fit SDMs: i) the Iberian bird species occurrences at 10km UTM square as a response or dependent variable; ii) the geographic layers of environmental information at 10km UTM square for the Iberian Peninsula as predictors or independent variables, such as climate data, ecosystem functioning attributes (EFAs) and the combined climate and EFA data. The dataset is provided by four <em>*.csv</em> files named as:</p> <p><em>1) The_Iberian_bird_species_occurrences_dataset_10km.csv</em></p> <p><em>2) CHELSA_bioclimate_variables_IP10km.csv</em></p> <p><em>3) MODIS_EVI-based_EFAs_IP10km.csv</em></p> <p><em>4) Combined_bioclimate_EFA_dataset_IP10km.csv</em></p> <p>For a more detailed description of the main dataset and each of these subdatasets, please refer to the attached README file.</p> <p><strong>Keywords:</strong> bird atlas, eBird data, ecosystem functional attributes (EFAs), Iberian Peninsula, IUCN categories, Model accuracy, MODIS EVI, narrow-ranged species, remote sensing, species distribution models (SDMs), widespread species</p>
Input features and benchmark data sets for protein complex prediction and E. coli proteome application by AF2Complex
<p>Benchmark data sets of AF2Complex, input features for application to E. coli proteome, and predicted structural models of E. coli Ccm I as described in</p> <p><strong>Predicting direct physical interactions in multimeric proteins with deep learning</strong></p> <p><em>Mu Gao, Davi Nakajima An, Jerry M. Parks, Jeffrey Skolnick</em></p> <ol> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/af2complex_bench.tar.gz">af2complex_bench.tar.gz</a>: Benchmark data sets CP17, Dimer1193 and Oligomer562, including input features for AF2Complex/AF-Multimer, both paired and unpaired MSAs, as well as sequences, experimental structures, and results presented in the AF2Complex work (~90GB de-compressed size)</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_Ccm_I.tar.gz">ecoli_Ccm_I.tar.gz</a>: Computational models of the<em> E. coli</em> Ccm I system</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_set.tar.gz">ecoli_sets.tar.gz</a>: Lists of benchmark sets of positive and negative PPIs from <em>E. coli</em></li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_af_fea.tar.gz">ecoli_af_fea.tar.gz: </a>Pre-generated input features of <em>E. coli</em> proteome for protein complex prediction and modeling by AF2Complex (4,429 proteins, ~800 GB de-compressed size). This data set can be used with AF2Complex to probe the interactions of any combinations among the 4,429 proteins of E. coli.</li> </ol> <p> </p>
AirGAM 2022r1 input data for all stations 2005-2019
<p>Contains EEA Airbase/AQ e-Reporting data and ECMWF ERA 5 met data for all stations.</p>
Dipolar Relaxation of Water Protons in the Vicinity of a Collagen-Like Peptide: Input Files for Simulation
<p>Input files to run the simulations in the paper</p> <p>Journal: The Journal of Physical Chemistry B<br> Title: Dipolar relaxation of water protons in the vicinity of a collagen-like peptide<br> Authors: Jouni Karjalainen, Henning Henschel, Mikko J. Nissi, Miika T. Nieminen, Matti Hanni<br> DOI: 10.1021/acs.jpcb.2c00052</p>
Inputs for lake morphology in ISIMIP3 runs of the global lake sector
<p>Empirical evidence demonstrates that lakes and reservoirs are warming across the globe. Consequently, there is an increased need to project future changes in lake thermal structure and resulting changes in lake biogeochemistry in order to plan for the likely impacts. Previous studies of the impacts of climate change on lakes have often relied on a single model forced with limited scenario-driven projections of future climate for a relatively small number of lakes. As a result, our understanding of the effects of climate change on lakes is fragmentary, based on scattered studies using different data sources and modelling protocols, and mainly focused on individual lakes or lake regions. This has precluded identification of the main impacts of climate change on lakes at global and regional scales and has likely contributed to the lack of lake water quality considerations in policy-relevant documents, such as the Assessment Reports of the Intergovernmental Panel on Climate Change (IPCC). The Lake Sector of the Inter-Sectoral Impact Model Intercomparison Project (ISIMIP) was founded for simulating climate change impacts on lakes using an ensemble of lake models and climate change scenarios. The protocol prescribes lake simulations driven by climate forcing from gridded observations and different Earth system models under various Representative Greenhouse Gas Concentration Pathways, all consistently bias-corrected on a 0.5° × 0.5° global grid. The ISIMIP Lake Sector is the largest international effort to project future water temperature, thermal structure, and ice phenology of lakes at local and global scales and paves the way for future simulations of the impacts of climate change on water quality and biogeochemistry in lakes.</p> <p>For a comprehensive description of ISIMIP, sectors, and protocols please see https://www.isimip.org/</p> <p>In order to simulate the impacts of climate change on lakes worldwide, the ISIMIP3 Lake Sector protocol has defined a set of lakes to be modelled by all participating lake models, as well as the basic morphometry (i.e., hypsographic curves) information for each lake. This repository includes all calculations performed to obtain the final set of lake location and mosphometry inpput information for ISIMIP3 runs. For this, available datasets on global lake extension and morphometry were used first for selecting a set of representative lakes on Earth (~40000 lakes, one for each 0.5º pixel of the normalized input/putput grid for ISIMIP across sectors), and then morphological characteristics were assigned to each representative lake using a database on global lake morphology. The final set of files constitute the input data for lake morphology and location for ISIMIP3 Lake Sector runs, which are produced in netCDF format.<br> </p>
Data from: Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range - wild grapevine sampling locations, Maxent input files, morphological and microsatellite data
<p><span>This dataset contains raw data described in the paper: "Rahimi O., Ohana-Levi N., Brauner H., Inbar N., Hübner S. and Drori E. (2021) "Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range", accepted for publication in "Ecology and Evolution".</span></p> <p><span>The spatial distribution of plants is constrained by demographic and eco-geographic factors that determine the range and abundance of the species. In this study, we performed genetic and morphological analyzes based on SSR and OIV datasets. In addition, according to the spatial distribution model performed by Maxent software we found that distance to water sources, Normalized difference vegetation index, and precipitation are the main environmental factors constraining <i>V.v. sylvestris</i> distribution at its southern distribution range. All raw data used for this study can be found in this deposit which contains a table with grapevine locations, Maxent input files, morphological and microsatellite data. </span></p>
Model Input Files
<p>The necessary input files to reproduce the model results of episodic and mobile convection with plume-slab termination. Files must be converted to .prm to work with ASPECT.</p>
Input files for publication: Binding of DEP Domain to Phospholipid Membranes: More Than Just Electrostatics
<p>Input files for publication: Francesco L. Falginella, Marek Kravec, Martina Drabinová, Petra Paclíková, Vítězslav Bryja, and Robert Vácha: Binding of DEP Domain to Phospholipid Membranes: More Than Just Electrostatics</p>
Input files and data for sensitivity analysis of ParFlow-CLM
<p>This zip file includes all the input files and forcing data for the sensitivity analysis of ParFlow-CLM.</p> <p>The file of "Stettbach_SA" includes all the information for the sensitivity analysis of Stettbach catchment.</p> <p>The file of "high_slope_SA" includes all the information for the sensitivity analysis of high slope effect in section 2.5.1 Topography effects.</p> <p>The file of "low_slope_SA" includes all the information for the sensitivity analysis of low slope effect in section 2.5.1 Topography effects.</p> <p>The file of "Barcelona_SA" includes all the information for the sensitivity analysis of Barcelona in section 2.5.2 Climate effects.</p> <p>The file of "LK_SA" includes all the information for the sensitivity analysis of Lemmenjoke national park in Finland in section 2.5.2 Climate effects.</p> <p>The file of "Stettbach_2020_2021" includes all the information for the runoff comparison between simulation and observation in 2020 and 2021 in section 3.1 (Model performance in Stettbach catchment).</p> <p>The file of "results_process_code" includes all the scripts used for the analysis of this study.</p>
Data and modeling input and parameter files for the 2021 Acapulco, Mexico earthquake and tsunami
<p>Raw and processed strong motion, GNSS, InSAR and tide gauge data for the event. Also includes MudPy slip inversion parameter files as well as GeoClaw input files.</p>
CNN-based Network Application for Petrophysical Parameter Inversion: Sensitivity Analysis of Input-output Parameters and Network Architecture
<p>The uploaded file includes four groups of data. They have original elastic and reservoir parameters data, the k=5 and k=25 (k means the sampling interval in inline and crossline.), and testing dataset.</p>
Marbled Murrelets prefer stratified waters close to freshwater inputs in Haida Gwaii, BC, Canada
<p>The Marbled Murrelet (Brachyramphus marmoratus) is a small seabird that is currently listed as threatened in Canada. Understanding this species' marine habitat preferences plays a vital role in our ability to focus conservation planning. We used the longest-running at-sea survey dataset available in British Columbia to examine hotspot persistence and habitat use at Laskeek Bay, Haida Gwaii, BC. The Laskeek Bay Conservation Society has been conducting spring and summer surveys along fixed transect routes in open and shoreline waters from 1997‒2018. Along with analyzing this long-term dataset, we conducted surveys to measure oceanographic variables (2018–2019) and tested whether murrelets in the same area used prey and oceanographic information to select marine habitat in conjunction with physical habitat features. Our hotspot persistence map, defined as areas that repeatedly had counts above a 75% threshold relative to other areas during a given survey, showed that murrelets consistently preferred shoreline transects. Murrelets also preferred shallow marine areas closer to streams, above higher proportions of sandy substrate and closer proximity to abundant nesting habitat. Modeling weather and time variables contributed little additional predictive power. Nonetheless, models that included physical environmental, oceanographic, and prey variables outperformed those with only physical environmental variables. Stratified water was the oceanographic variable most strongly related to higher counts. Our study suggests that stratified waters could work with stream systems to create productive zones for foraging murrelets, and highlights the importance of murrelets having access to marine areas with the preferred physical features. </p>
Example of input and output datasets of the cost module implemented in the SHERPA modelling tool
<p><em>Integrated Assessment Model provides a useful framework for evaluating different aspects of air quality policies, spanning from abatement measures to emissions, concentrations, health impacts and costs. These models are then useful to provide a holistic view of the impacts of policies, so that ex-ante one can evaluate how various policies will impact air concentrations, health benefits and implementation costs. Among these Integrated Assessment Models, SHERPA (Screening for High Emission Potentials on Air) has been recently used to evaluate the impact of policies, covering all aspects from measures to health, but without being able to provide the dimension related to abatement measures costs. These three files dataset provide an example of the ability of the new cost module in SHERPA to calculate the cost of emission reductions.</em></p>
Sample Input Data and Supporting Files for the SELECT Model of Urbanization
<p>Sample Input Data and Supporting Files for the SELECT Model of Urbanization</p> <p>Code available at: https://github.com/IMMM-SFA/select</p>
AURES II, WP8, Dataset Input Data
<p>Input Dataset for power system modelling with the model BALMOREL conducted in the course of the AURES II project.<br> Workpackage 8; TU Wien</p> <p> </p>
Input files for "Predicting the orientation of adsorbed proteins steered with electric fields using a simple electrostatic model"
<p>In files to run PyGBe (https://github.com/pygbe/pygbe), and reproduce the results from the J. Phys. Chem. B article "Predicting the orientation of adsorbed proteins steered with electric fields using a simple electrostatic model". </p>
Artifacts for the 2022 ATVA Paper: SISL: Concolic Testing of Structured Binary Input Formats via Partial Specification
<p>Artifacts for the evaluation of the publication <em>SISL: Concolic Testing of Structured Binary Input Formats via Partial Specification,</em> which will be published as part of the 2022 <em>International Symposium on Automated Technology for Verification and Analysis </em>(<a href="https://atva-conference.org/2022/">ATVA</a>).</p> <p>Usage of the provided artifacts requires installing the SISL tooling first. In this regard, please refer to either the <a href="https://agra-uni-bremen.github.io/sisl/">SISL user manual</a> or the documentation in the <em><a href="https://github.com/agra-uni-bremen/sisl">sisl</a> </em>and <a href="https://github.com/agra-uni-bremen/sisl-vp"><em>sisl-vp</em></a> GitHub repositories. Both <em>sisl</em> and <em>sisl-vp</em> need to be installed to system-wide locations in order to be able to use the scripts provided as part of these artifacts.</p> <p>More information on using the artifacts is available in the README.md file.</p>
Data from: Between the Cape Fold Mountains and the deep blue sea: comparative phylogeography of selected codistributed ectotherms reveals asynchronous cladogenesis. Sampling locations and MaxEnt input files
<p><span>We compare the phylogeographic structure of thirteen codistributed ectotherms including four reptiles (a snake, a legless skink and two tortoise species) and nine invertebrates (six freshwater crabs and three velvet worm species) to test the presence of congruent evolutionary histories. </span><span>Phylogenies were estimated and dated using maximum likelihood and Bayesian methods with combined mitochondrial and nuclear DNA sequence datasets. </span><span>All taxa demonstrated a marked east/west phylogeographic division, separated by the Cape Fold Mountain range.</span><span> <span>Phylogeographic concordance factors were calculated to assess the degree of evolutionary congruence among the study species and </span></span><span>generally supported a shared pattern of diversification along the east/west longitudinal axis</span><span>. Testing simultaneous divergence between the eastern and western phylogeographic regions indicated </span><span>pseudo-congruent evolutionary histories among the study taxa, with at least three separate divergence events throughout the Mio/Plio/Pleistocene epochs.</span><span> <span>Climatic refugia were identified for each species using climatic niche modeling, </span></span><span>demonstrating taxon-specific responses to climatic fluctuations. Climate and the Cape Fold Mountain barrier explained the highest proportion of genetic diversity in all taxa, while climate was the most significant individual abiotic variable. </span><span>This study highlights the complex interactions between the Cape Fold Mountains and past climatic oscillations during the Mio/Plio/Pleistocene. The congruent east/west phylogeographic division observed in all taxa lends support to the conclusion that the longitudinal climatic gradient within the Greater Cape Floristic Region, mediated in part by the barrier to dispersal posed by the Cape Fold Mountains, plays a major role in lineage diversification and population differentiation.</span></p>
CeSta_makeFeatures_input
<p>This tarball contains all input data needed to generate engineered multi-omic features which are taken as input by our CeSta Cetuximab sensitivity classifier.</p>
Inputs and results of "A quantitative and qualitative citation analysis to retracted articles in the humanities domain"
<p>This repository contains the datasets and visualizations generated in our work: <strong>"A quantitative and qualitative citation analysis to retracted articles in the humanities domain"</strong>.</p> <p><strong>Note:</strong> the data are all contained inside the <strong><em>data.zip</em> </strong>file. You need to unzip the container to get access to all the files and directories listed below.</p> <p>The data (citations) gathered accompanied by their annotated characteristics are stored in <strong><em>data/</em>:</strong></p> <ul> <li><em>cits.csv: </em>a dataset containing all the entities (rows in the CSV) which have cited a retracted article in the humanities domain. Each citing entity (row) is accompanied by a set of features (columns) that characterizes it.<br> <strong>Note: </strong>this dataset is licensed under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.</li> <li><em>content.csv: </em>a dataset containing the abstracts and the in-text citation contexts of all the citing entities gathered.<br> <strong>Note: </strong>the data keep their original license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> <li><em>excluded_hum_retractions.csv: </em>a list of the 12 humanities retracted articles with a humanities affinity score < 2, therefore excluded from the analysis. </li> </ul> <p> </p> <p><strong>Topic modeling</strong></p> <p>We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the <em><strong>topic_model/</strong></em> directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools and creating a completely customizable visual workflow [1]. The directory <em><strong>workflow/ </strong></em>contains the workflows used in MITAO. The topic modeling results for each textual feature are separated into two different folders, <em><strong>abstract/</strong></em> for the abstracts, and <em><strong>cits_context/</strong></em> for the in-text citation contexts. Both the directories contain the following directories/files: </p> <ul> <li> <p><em><strong>datasets_and_views/: </strong></em>the datasets and visualizations generated using MITAO. </p> </li> <li> <p><em><strong>ldamodel_corpus_dict/: </strong></em>it contains the dictionary, the LDA topic model, and the tokenized and vectorized corpus.</p> </li> <li><em><strong>rawdata/: </strong></em>the textual collection, metadata, and stopwords used as input in the workflow of MITAO</li> </ul> <p> </p> <p><strong>References</strong></p> <p>[1] Ferri, P., Heibi, I., Pareschi, L., & Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135–149. <a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a></p> <p> </p> <ol> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.