Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
608
datasets available to search
ShareScore release 0.9.0
Dataset results
608 results for “ensembles”
Figure 1 in Ensemble distribution modeling of the Mesopotamian spiny-tailed lizard, Saara loricata (Blanford, 1874), in Iran: an insight into the impact of climate change
Figure 1. The presence records (black dots) used for the development of a maximum entropy model for predicting the habitat suitability of the Mesopotamian spiny-tailed lizard.
Figure 4 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 4. Mean suitability for Arborophila crudigularis under baseline climate conditions and future climate scenarios. cccma and csiro represent two general circulation models; RCP2.6, and RCP8.5 represent two greenhouse gas emission scenarios; EN is entire suitable habitat; PR is presence records.
Figure 2 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 2. Performance of each model for predicting the suitable habitat for Arborophila crudigularis. GLM: Generalized linear model; GBM: generalized boosting model; GAM: generalized additive model; CTA: classification tree analysis; ANN: artificial neural network; FDA: flexible discriminant analysis; MARS: multiple adaptive regression splines; RF: random forest; MAXENT: maximum entropy model.
Figure 6 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 6. Changes in suitable habitat for Arborophila crudigularis under the RCP8.5 emission scenario. cccma and csiro represent two general circulation models.
Figure 5 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 5. Changes in suitable habitat for Arborophila crudigularis under the RCP2.6 emission scenario. cccma and csiro represent two general circulation models.
Ensemble Learning of Catchment-Wise Optimized LSTMs Enhances Regional Rainfall-Runoff modelling - Case Study: Basque Country, Spain - Data
<div> <div>This data and results are for paper: "Ensemble Learning of Catchment-Wise Optimized LSTMs Enhances Regional Rainfall-Runoff modelling - Case Study: Basque Country, Spain" by Hosseini et al. 2024 (Preprint - Under review J.Hydro 2024) Available at SSRN: <a href="https://ssrn.com/abstract=4918782" target="_blank" rel="noopener">https://ssrn.com/abstract=4918782</a></div> <div> </div> </div>
Changes in four decades of near-CONUS tropical cyclones in an ensemble of 12km thermodynamic global warming simulations
<p>Snapshot level data of TC extractions from the thermodynamical global warming runs described in "Changes in four decades of near-CONUS tropical cyclones in an ensemble of 12km thermodynamic global warming simulations."</p>
Supplementary Dataset: Air quality modeling intercomparison and multi-scale ensemble chain for Latin America
<p>The Supplementary dataset of the manuscript titled "Air quality modeling intercomparison and multi-scale ensemble chain for Latin America" can be downloaded via this link:<br>https://swiftbrowser.dkrz.de/public/dkrz_3ab03fbe-db0a-42e8-8b19-caf61d10634d/PAPILA/</p> <p>The data repository contains the model data used in the model intercomparison with six global and regional chemical-transport model over Latin America and the observation datasets used in the model intercomparison. This work presents the first model intercomparison and ensemble construction for Latin America, which was assembled under the Prediction of Air Pollutants in Latin America (PAPILA project (https://papila-h2020.eu/papila). </p>
Data and code for FishMIP global marine ecosystem model ensemble projections summarised by countries and territories and other selected marine spatial regions.
<p>R code to extract and create data tables and summary plots of FishMIP mean ensemble projections provided are for percentage change in "exploitable fish biomass", which is a proxy for the biomass available to fisheries, consisting of marine animals spanning the size range 10 g to 100 kg: this is typically dominated by fish, but is also inclusive of other animals such as crustaceans and cephalopods.</p> <p>This release contains scripts and summary data for producing figures in Part A of the following report:</p> <p>Blanchard, J.L., Novaglio, C., eds. (2024). Climate change risks to marine ecosystems and fisheries: Future projections from the Fisheries and Marine Ecosystems Model Intercomparison Project. FAO Fisheries and Aquaculture Technical Paper No. 707. Rome, FAO.</p> <p>Please refer to the above report to cite and for more information.</p> <p>The summary data are here:</p> <p>https://github.com/Fish-MIP/FAO_Report/blob/main/data/table_stats_formatted_admin_full.csv</p> <p>Where the column 'spatial_scale' refers to the type of aggregation:</p> <p>FAO_area = High Sea areas grouped by FAO Major Fishing Areas</p> <p>countries = Exclusive Economic Zones</p> <p>countries_admin = Exclusive Economic Zones results aggregated into Administrative Countries</p> <p>Please note that these results can also be visualised and downloaded from our shiny app: https://rstudio.global-ecosystem-model.cloud.edu.au/shiny/FAO_report_shiny/</p> <p> </p>
North Atlantic average sea-surface temperature in a CMIP6 multi-model ensemble of historical and ssp585/ssp245 simulations
<p>This dataset contains North Atlantic average sea-surface temperatures and derived indices calculated for a multi-model ensemble of historical and future scenario (ssp585, ssp245) simulations contributing to the Coupled Model Intercomparison Project phase 6. A detailed description of the dataset is provided by Zanchettin, D., and Rubino, A., Accelerated North Atlantic surface warming reshapes the Atlantic Multidecadal Variability, Communications Earth & Environment, 2024, doi:10.1038/s43247-024-01804-x.</p> <p><br>The data are provided as netcdf files.</p> <p>The name of each file is structured as {model}_r{realization}_historical_{scenario}.nc where {model} is the model name, {realization} is a number corresponding to the historical realization, and {scenario} is either of the two scenarios considered (ssp585 and ssp245).</p> <p>Each file contains data for the following one dimensional variables:</p> <ul> <li>year: the sequence of years for which the data are provided</li> <li>NASST: annual-average spatially averaged North Atlantic sea-surface temperature</li> <li>state: slowly variable component of NASST obtained from a dlm decomposition of NASST</li> <li>strend: stochastic trend of NASST obtained from a dlm decomposition of NASST</li> <li>AMV: Atlantic Multidecadal Varibility index obtained as difference between NASST and state</li> </ul> <p> </p>
AlphaFold2-Based Characterization of Apo and Holo Protein Structures and Conformational Ensembles Using Randomized Alanine Sequence Scanning Adaptation: Capturing Shared Signature Dynamics and Ligand-Induced Conformational Changes
<p>Proteins often exist in multiple conformational states, influenced by the binding of ligands or substrates. The study of these states, particularly the apo (unbound) and holo (ligand-bound) forms, is crucial for understanding protein function, dynamics, and interactions. In the current study, we use AlphaFold2 that combines<span> randomized</span> <span><span> </span>alanine<span> </span>sequence masking<span> </span>with shallow multiple sequence alignment<span> </span>subsampling to expand the conformational diversity of the predicted structural<span> </span>ensembles and<span> </span>capture conformational changes between apo and holo protein forms. Using several well-established datasets of<span> </span>structurally diverse apo-holo protein pairs, the proposed approach </span><span>enables<span> </span>robust predictions of apo and holo structures and conformational ensembles, while also displaying notably similar dynamics distributions. These observations are consistent with<span> </span>the view </span><span> </span>that the intrinsic dynamics of allosteric proteins is defined by the structural topology of the fold and favors conserved conformational motions driven by soft modes among orthologs. We also found<span> </span>a significant <span>correlation </span>between conformational flexibility and <span> </span>AlphaFold2 metric of statistical significance pLDDT for the apo-holo pairs in which ligand binding induced local moderate conformational changes. For apo-holo pairs exhibiting larger structural changes, this relationship<span> </span>becomes nonlinear, reflecting inability of AlphaFold2 confidence metrics to identify high energy functional conformations. Our findings support the notion that AlphaFold2 approaches can yield reasonable accuracy in predicting minor conformational adjustments between apo and holo states, especially for proteins with <span> </span>moderate localized changes upon ligand binding. However, for large, hinge-like domain movements, AF2 tends to predict the most stable domain orientation which is typically the apo form rather than the full range of functional conformations characteristic of the holo ensemble. These results indicate that modeling of multiple functional states of proteins may require more accurate detection of flexible region conformations and cannot solely rely on the pLDDT metric as the major determinant of the prediction accuracy in reproducing functional conformational ensembles.<span> </span></p>
ECOCLIMAP-SG-ML: an ensemble land cover map for numerical weather prediction
<p>This dataset contains ensemble land cover maps for numerical weather prediction at 60 m resolution over Europe. As they were<br>generated thanks to machine learning, the weights and the training data are also provided.</p>
A large-ensemble simulation of yields and meteorological drivers to evaluate spatial compounding crop failures in Europe
<p>The dataset consists of a subset from Vogel et al. (2021), comprising large-ensemble simulations of winter wheat yields aggregated at the country level for 20 European countries. The winter wheat yields were simulated by the APSIM-Wheat model (version 7.10) (Zhang et al. 2014) driven by meteorological data from the EC-Earth global climate model (Hazeleger et al., 2010; Van der Wiel et al., 2019). To investigate meteorological drivers of crop failure, the dataset also includes monthly means of daily precipitation, vapour pressure deficit, and maximum temperature fields, of two leading European producers, i.e. France and Germany. For more details, see the description in Vogel et al. (2021).</p> <p>By using this data, you also agree to cite the reference below:</p> <p>Vogel, J., Rivoire, P., Deidda, C., Rahimi, L., Sauter, C. A., Tschumi, E., van der Wiel, K., Zhang, T., Zscheischler, J. (2021). Identifying meteorological drivers of extreme impacts: an application to simulated crop yields. Earth System Dynamics,12(1),151-172.</p>
Outputs from Isca perturbed parameter ensemble (PPE) simulations (Part I)
<p>Outputs from Isca perturbed paramter ensemble (PPE) simulations under 1xCO2 (control run) and 4xCO2 (perturbed run), in which the simulations are prescribed with Q-flux.</p> <ul> <li>cld_fbk_cmp.zip: The outputs used for the comparison of cloud feedback computation methods.</li> <li>qflux_clisccp_data.zip: The outputs contain the variable clisccp from control and perturbed runs, used for the cloud feedabck calculation</li> <li>qflux_extracted_data_toa_flux_30yr.zip: The last 30-year top of the atmosphere (TOA) flux data used for deriving the equilibrium climate sensitivity, total climate feedback and effective radiative forcing through the Gregory plot</li> <li>qflux_extracted_data.zip: The last 5-year simulation outputs for control and perturbed runs.</li> <li>qflux_extracted_data_10yr.zip: The last 10-year simulation outputs for control and perturbed runs.</li> </ul> <p>This is Part I, and Part II can be found at: <a href="https://doi.org/10.5281/zenodo.5188175">10.5281/zenodo.5188175</a></p>
Synthetic large ensembles from four observation-based products of sea-air CO2 flux from Olivarez et al. (2021)
<p>Synthetic large ensembles from four observation-based products of sea-air CO2 flux from Olivarez et al. (2021). Observation-based products are:</p> <p>Council for Scientific and Industrial Research-Machine Learning (CSIR-ML6)<br> Max Planck Institute Self-Organizing Map-Feed-Forward Neural Network (MPI-SOMFFN)<br> Jena, Germany-Max Planck Institute for Biogeochemistry-Mixed Layer Scheme (JENA-MLS)<br> Copernicus Marine Environment Monitoring Service Feed-Forward Neural Network (CMEMS-FFNN)</p>
Dataset for "A stacking ensemble algorithm for improving the biases of forest aboveground biomass estimations from multiple remotely sensed datasets"
<p>This dataset is associated with a research article entitled "A stacking ensemble algorithm for improving the biases of forest aboveground biomass estimations from multiple remotely sensed datasets".</p>
Datasets for "Needle in a Bayes Stack: a Hierarchical Bayesian Method for Constraining the Neutron Star Equation of State with an Ensemble of Binary Neutron Star Post-merger Remnants"
<p>All data used for "Needle in a Bayes Stack: a Hierarchical Bayesian Method for Constraining the Neutron Star Equation of State with an Ensemble of Binary Neutron Star Post-merger Remnants", Criswell, A.W., et al. (2022). The code used to create the paper results from this data can be found at <a href="https://github.com/criswellalexander/hbpm_paper">https://github.com/criswellalexander/hbpm_paper</a> and the underlying software package can be found at <a href="https://github.com/criswellalexander/bayestack">https://github.com/criswellalexander/bayestack</a>.</p>
Accurate Prediction of Enzyme Thermostabilization with Rosetta using AlphaFold Ensembles
<p>DT<sub>M</sub> vs DG<sub>f,mut</sub> values for scoring LovD, LipA, <em>p</em>-nitrobenzyl esterase, xylanase A and tryptophan 6-halogenase variants (<em>DTM_vs_DDGf_mut.xlsx</em>).</p> <p>AlphaFold predicted structures in PDB and Pymol sessions formats for top scoring LovD, LovD6, LovD9, LipA WT, LipA 6B, <em>p</em>-nitrobenzyl esterase WT, xylanase A WT and tryptophan 6-halogenase WT decoys (<em>mAF-min_ensembles.zip</em>).</p> <p>Rosetta energies for all calculations (<em>Rosetta_scores.zip</em>).</p>
Gene/Protein BridgeDb ID Mapping Database (Ensembl Metazoa 52)
<p>Mapping databases derived from Ensembl Metazoa 52. These files can be used with BridgeDb.<br> The scripts which were used to create these databases based on Ensembl BioMart can be found at <a href="https://github.com/bridgedb/create-bridgedb-genedb">https://github.com/bridgedb/create-bridgedb-genedb</a>.</p> <p>This work was funded by the <a href="https://fairplus-project.eu/">FAIRplus project</a> (grant agreement no 802750) and <a href="https://www.nwo.nl/en/researchprogrammes/open-science/open-science-fund/open-science-fund-2021-awarded-grants">NWO Open Science Fund</a> (grant no <a href="https://www.nwo.nl/en/projects/203001121">203.001.121</a>).</p>
Predictors and predictands for "Downscaling CORDEX through deep learning to daily 1 km multivariate ensemble in complex terrain"
<p>Predictors and predictands for "Downscaling CORDEX through deep learning to daily 1 km multivariate ensemble in complex terrain". Training predictors from the ERA5 reanalysis, projecting predictors from CORDEX EUR11, and predictand from ReKIS (https://rekis.hydro.tu-dresden.de). Data is saved in ".rds" format, to be read from R, except for CORDEX files in NetCDF.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.