Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

608

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

608 results for “ensembles”

Learn how ShareScore rates datasets ↗
zenodo44/100

An ensemble of trend preserving statistically downscaled projections for key marine variables under three different future scenarios for the Mediterranean Sea

<p>The ensemble provides future projections of key marine variables under climate change for the Mediterranean region. The datasets were produced for three different future scenarios (SSP1-2.6, SSP2-4.5 and SSP5-8.5) and five different variables (potential temperature, salinity, dissolved oxygen, pH and chlorophyll) at three different depth levels (5m, 25m and seafloor with the exception of chlorophyll) at monthly frequency for the years 1993 - 2099. The statistical metrics provided are the mean, standard deviation, minimum, maximum median, 2.5 and 97.5 percentile. The ensemble is computed over 3-7 different CMIP6 model realisations (depending on variable), the bias corrections and statistical downscaling was trained on the GLORYS12V1 reanalysis provided by the Copernicus Marine Environment Monitoring Service (CMEMS).</p> <p>The following Earth System Models were used in building the ensemble:</p> <ul> <li>CMCC-ESM2 (Lovato et al. 2022)</li> <li>CMCC-CM2-SR5 (Cherchi et al. 2019)</li> <li>GFDL-ESM4 (Dunne et al., 2020)</li> <li>MPI-ESM1-2-LR (Mauritsen et al., 2020)</li> <li>IPSL-CM6A-LR&nbsp;(Boucher et al. 2020)</li> </ul> <p>A description of the downscaling approach and evaluation of the datasets over the European regions is published in&nbsp;<a href="https://doi.org/10.1038/s41598-024-51160-1">Kristiansen et al. 2024</a>.</p> <p>&nbsp;<br>Analogue datasets are provided in separate zenodo entries for the regions of the North Sea, the Baltic Sea, the Bay of Biscay, the Chilean coast and the area around the Yucat&aacute;n Peninsula, see &ldquo;Related identifiers&rdquo;.</p> <p>We acknowledge the World Climate Research Programme's Working Group on Coupled Modelling, which is responsible for CMIP. Generated using E.U. Copernicus Marine Service Information; <a href="https://doi.org/10.48670/moi-00021">https://doi.org/10.48670/moi-00021</a>,&nbsp; <a href="https://doi.org/10.48670/moi-00019">https://doi.org/10.48670/moi-00019</a>.</p> <p><br><strong>This data is distributed under <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License</a>.</strong></p> <p>&nbsp;</p>

openother-ncMay 2022View details →
zenodo44/100

An ensemble of trend preserving statistically downscaled projections for key marine variables under three different future scenarios for the Baltic Sea

<p>The ensemble provides future projections of key marine variables under climate change for the Baltci Sea region. The datasets were produced for three different future scenarios (SSP1-2.6, SSP2-4.5 and SSP5-8.5) and five different variables (potential temperature, salinity, dissolved oxygen, pH and chlorophyll) at three different depth levels (5m, 25m and seafloor with the exception of chlorophyll) at monthly frequency for the years 1993 - 2099. The statistical metrics provided are the mean, standard deviation, minimum, maximum median, 2.5 and 97.5 percentile. The ensemble is computed over 3-7 different CMIP6 model realisations (depending on variable), the bias corrections and statistical downscaling was trained on the GLORYS12V1 reanalysis provided by the Copernicus Marine Environment Monitoring Service (CMEMS).</p> <p>The following Earth System Models were used in building the ensemble:</p> <ul> <li>CMCC-ESM2 (Lovato et al. 2022)</li> <li>CMCC-CM2-SR5 (Cherchi et al. 2019)</li> <li>GFDL-ESM4 (Dunne et al., 2020)</li> <li>MPI-ESM1-2-LR (Mauritsen et al., 2020)</li> <li>IPSL-CM6A-LR&nbsp;(Boucher et al. 2020)</li> </ul> <p>A description of the downscaling approach and evaluation of the datasets over the European regions is published in&nbsp;<a href="https://doi.org/10.1038/s41598-024-51160-1">Kristiansen et al. 2024</a>.</p> <p>Analogue datasets are provided in separate zenodo entries for the regions of the Mediterranean Sea, the North Sea, the Bay of Biscay, the Chilean coast and the area around the Yucat&aacute;n Peninsula, see &ldquo;Related identifiers&rdquo;.</p> <p>We acknowledge the World Climate Research Programme's Working Group on Coupled Modelling, which is responsible for CMIP. Generated using E.U. Copernicus Marine Service Information; <a href="https://doi.org/10.48670/moi-00021">https://doi.org/10.48670/moi-00021</a>,&nbsp; <a href="https://doi.org/10.48670/moi-00019">https://doi.org/10.48670/moi-00019</a>.</p> <p><br><strong>This data is distributed under <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License</a>.</strong></p>

openother-ncMay 2022View details →
zenodo44/100

Biased ensembles of pulsating active matter: figure data

<p>The following is a zip file containing figure data for all figures published in the manuscript tittled 'Biased ensembles of pulsating active matter', available in archive: https://arxiv.org/abs/2403.16961</p> <p>v2 includes the updated data for figure 2. Otherwise all remain as before.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Ensemble statistics for modelled Eddy Kinetic Energy in the Southern Ocean

<p>This dataset contains surface eddy kinetic energy over the Southern Ocean region, sourced from a 50-member ensemble of 0.25&deg; ocean model simulations. It is used in the paper &quot;Circumpolar variations in the chaotic nature of Southern Ocean eddy dynamics&quot; published in Journal of Geophysical Research - Oceans.</p> <p>This dataset has been computed from the OceaniC Chaos &ndash; ImPacts, strUcture, predicTability (OCCIPUT) global ocean/sea-ice ensemble simulation. It is composed of 50 members with a horizontal resolution of 1/4&deg; and 75 geopotential levels (<a href="http://doi.org/10.5194/gmd-10-1091-2017">Bessi&egrave;res et al., 2017</a>, Penduff et al., 2014). The numerical configuration is based on the version 3.5 of the NEMO model (<a href="https://www.nemo-ocean.eu/doc">Madec, 2008</a>). The 50 members were started on January 1st 1960 from a common 21-year spinup. A small stochastic perturbation is applied to the equation of state of sea water (as in <a href="https://doi.org/10.1016/j.ocemod.2013.02.004">Brankart, 2013</a>) within each member during 1960, then switched off during the rest of the simulation. This 1-year perturbation generates an ensemble spread which grows and saturates after a few months up to a few years depending on the region. The 50 members are driven through bulk formulae during the whole 1960-2015 simulation by the same realistic 6-hourly atmospheric forcing (Drakkar Forcing Set DFS5.2, Dussin et al., 2016) derived from ERA interim atmospheric reanalysis. Data is for the period 1979-2015.</p> <p>The sea level anomaly is found according to <a href="http://doi.org/10.1016/j.pocean.2020.102314">Close et al (2020)</a> and converted into surface geostrophic velocity anomaly using the geostrophic relation. This velocity field is then used to calculate the eddy kinetic energy (EKE). Data is averaged over calendar month, and restricted to the latitude range 40&deg;-60&deg;S. A full description of this process is included in the companion paper.</p> <p>The dataset includes EKE files (eke_0??.nc), with monthy EKE saved for the period 1979-2015 for each ensemble member, and a single file (tau.nc) for the monthly-averaged wind stress over the same period.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Second release of the data associated with the paper entitled 'Cluster-enhanced ensemble learning for mapping global monthly surface ozone from 2003 to 2019'

<p>This is the second release of the data associated with the paper entitled &#39;Cluster-enhanced ensemble learning for mapping global monthly surface ozone from 2003 to 2019&#39;.</p> <p>The paper was published&nbsp;in&nbsp;Geophysical Research Letters. We provide the data that has been smoothed by moving filter&nbsp;and not. The data can be loaded by the <em>raster </em>package in <em>R.</em>&nbsp;Note that the unit is ppmv.</p> <p>Please note that both of these files must be in the same directory to open in <em>R</em> properly<em>.</em></p> <p>Please get in touch with the authors if you have any issues, email: xliu21@smail.nju.edu.cn or wanghk@nju.edu.cn</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Gene/Protein BridgeDb ID Mapping Database (Ensembl Fungi 49)

<p>Ensembl Fungi 49 derived ID mapping database for use with BridgeDb.<br> The&nbsp;scripts used to create these databases based on Ensembl BioMart&nbsp;can be found at <a href="https://github.com/bridgedb/create-bridgedb-genedb">https://github.com/bridgedb/create-bridgedb-genedb</a>.</p> <p>This work was funded by the&nbsp;<a href="https://fairplus-project.eu/">FAIRplus project</a>&nbsp;(grant&nbsp;agreement no 802750) and&nbsp;<a href="https://www.nwo.nl/en/researchprogrammes/open-science/open-science-fund/open-science-fund-2021-awarded-grants">NWO Open Science Fund</a>&nbsp;(grant no&nbsp;<a href="https://www.nwo.nl/en/projects/203001121">203.001.121</a>).</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

EPTGODD-WHU: Ensemble Precipitation and Temperature from CMIP6 GCMs optimized by OLS-DT-DNN methods integration (1850-2100)

<p>This monthly global climate dataset EPTGODD-WHU (precipitation and mean temperature variables with grid size of 0.5&deg;&times;0.5&deg;) was ensembled from 16 selected CMIP6 GCMs. The published dataset was optimized by OLS (Ordinary Linear Square)-DT (Decision Tree)-DNN (Deep Neural Network) methods integration. The CF (Climate and Forecast) v1.6 was employed as the guideline for NetCDF4 format. The periods of temperature files can be divided into historical (1850-1900) and future (2015-2100) periods. For precipitation, this product provides future (2015-2100) period. Three future scenarios (SSP1-2.6, SSP2-4.5 and SSP5-8.5) were selected for both variables. The units of this dataset are degrees Celsius and mm/month for temperature and precipitation, respectively. Each NetCDF4 file in this dataset includes three dimensions (time, latitude (-89.75&deg;N to 89.75&deg;N) and longitude (-179.75&deg;E to 179.75&deg;E)).</p>

opencc-by-4.0May 2022View details →
zenodo44/100

R-code for publication: Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods

<p>This is the R-code as well as the underlying data&nbsp;needed to reproduce the results of the springer book chapter: &quot;Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods&quot;</p> <p>For more information contact:&nbsp;dlieske@mta.ca</p> <p>&nbsp;</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

CREMP-CycPeptMPDB: Conformer-rotamer ensembles of macrocyclic peptides for machine learning with permeability annotations

<p>CREMP-CycPeptMPDB: A resource generated for the rapid development and evaluation of machine learning models for permeable macrocyclic peptides. CREMP-CycPeptMPDB contains 3,258 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 8.7 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations and with experimental membrane permeability measurements obtained from the <a href="http://cycpeptmpdb.com/" target="_blank" rel="noopener">CycPeptMPDB</a> database. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>This dataset complements the <a title="CREMP" href="../doi/10.5281/zenodo.7931444" target="_blank" rel="noopener">CREMP dataset</a>, which contains a larger selection of conformer ensembles for homodetic macrocyclic peptides.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the &ldquo;pickle&rdquo; folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single &ldquo;pickle.tar.gz&rdquo; archive. In the &ldquo;sdf_and_json&rdquo; folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, &ldquo;sdf_and_json.tar.bz2&rdquo;. A single summary CSV file is also provided containing &rdquo;sequence&rdquo;, &ldquo;smiles&rdquo;, &ldquo;num_monomers&rdquo;, &ldquo;num_atoms&rdquo;, &ldquo;num_heavy_atoms&rdquo;, along with the CREST metadata &ldquo;totalconfs&rdquo;, &ldquo;uniqueconfs&rdquo;, &ldquo;lowestenergy&rdquo;, &ldquo;poplowestpct&rdquo;, &ldquo;temperature&rdquo;, &ldquo;ensembleenergy&rdquo;, &ldquo;ensembleentropy&rdquo;, and &ldquo;ensemblefreeenergy&rdquo;. The number of unique conformers with different 3D structures is given by &ldquo;uniqueconfs&rdquo;, while &ldquo;totalconfs&rdquo; includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 13 GB for "pickle.tar.gz" and 84 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Data of the publication Hot ion implantation to create dense NV center ensembles in diamond

<p>Data of the publication published under the reference:&nbsp;M.W.<em> </em>Ngambeu Ngambou, Appl. Phys. Lett. 124, 134002 (2024).</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

PPMLES – Perturbed-Parameter ensemble of MUST Large-Eddy Simulations

<h2>Dataset description</h2> <p>This repository contains the PPMLES (Perturbed-Parameter ensemble of MUST Large-Eddy Simulations) dataset, which corresponds to the main outputs of 200 large-eddy simulations (LES) of microscale pollutant dispersion that replicate the MUST field experiment [Biltoft. 2001, Yee and Biltoft. 2004] for varying meteorological forcing parameters.</p> <p>The goal of the PPMLES dataset is to provide a comprehensive dataset to better understand the complex interactions between the atmospheric boundary layer (ABL), the urban environment, and pollutant dispersion. It was originally used to assess the impact of the meteorological uncertainty on microscale pollutant prediction and to build a surrogate model that can replace the costly LES model [Lumet et al. 2025]. The total computational cost of the PPMLES dataset is estimated to be about 6 million core hours.</p> <p>For each sample of meteorological forcing parameters (inlet wind direction and friction velocity), the <a href="https://www.cerfacs.fr/avbp7x/">AVBP</a> solver code [Schonfeld and Rudgyard. 1999, Gicquel et al. 2011] was used to perform LES at very high spatio-temporal resolution (1e-3s time step, 30cm discretization length) to provide a fine representation of the pollutant concentration and wind velocity statistics within the urban-like canopy. The total computational cost of the PPMLES dataset is estimated to be about 6 million core hours.</p> <h2>File list</h2> <p>The data is stored in <a href="https://www.hdfgroup.org/solutions/hdf5/">HDF5</a> files, which can be efficiently processed in Python using the <a href="https://docs.h5py.org/en/stable/index.html">h5py</a> module.&nbsp;</p> <ul> <li><em>input_parameters.h5:</em> list of the 200 input parameter&nbsp;samples<em>&nbsp;(alpha_inlet, ustar) </em>obtained using the&nbsp;Halton sequence that defines the PPMLES ensemble.</li> <li><em>ave_fields.h5</em>: lists of the main field statistics predicted by each of the 200 LES samples over the 200-s reference window [Yee and Biltoft. 2004], including: <ul> <li><em>c:</em> the time-averaged pollutant concentration in ppmv <em>(dim = (n_samples, n_nodes) = (200, 1878585))</em>,&nbsp;</li> <li><em>(u, v, w):&nbsp;</em>the time-averaged wind velocity components in m/s,</li> <li><em>crms: </em>the root mean square concentration fluctuations in ppmv,&nbsp;</li> <li><em>tke:</em> the turbulent kinetic energy in m^2/s^2,</li> <li><em>(uprim_cprim, vprim_cprim, wprim_cprim)</em>: the pollutant turbulent transport components</li> </ul> </li> <li><em>uncertainty.h5</em>: lists of&nbsp;the estimated aleatory uncertainty induced by the internal variability of the LES&nbsp;<em>(variability_#)</em> [Lumet et al. 2024] for each of the fields in <em>ave_fields.h5</em>. Also includes the stationary bootstrap [Politis and Romano. 1994] parameters <em>(n_replicates, block_length) </em>used to estimate the uncertainty for each field and each sample.</li> <li><em>mesh.h5</em>: the tetrahedral mesh on which the fields are discretized, composed of about 1.8 millions of nodes.</li> <li><em>time_series.h5</em>:&nbsp;HDF5 file consisting of 200 groups (<em>Sample_NNN</em>) each containing the time series of the pollutant concentration (c) and wind velocity components (u, v, w) predicted by the LES sample #NNN at 93 locations.&nbsp;</li> <li><em>probe_network.dat</em>: provides the location of each of the 93 probes corresponding to the positions of the experimental campaign sensors [Biltoft. 2001].</li> </ul> <p><strong>Warning:</strong> the propylene concentration are expressed in ppmv, except in time_series.h5 in which they are given as mass fractions. To convert them in ppmv, the formula is: <code>c = c * (rho/rho_propylene) * 10**6</code> with (rho/rho_propylene) = 0.66 the density ratio between air and propylene.</p> <h2>Code examples</h2> <p>In the following, examples of how to use the PPMLES dataset in Python are provided. These examples have the following dependencies:&nbsp;</p> <pre><code>requires-python = "&gt;=3.9" dependencies = [ "h5py==3.8.0", "numpy==1.26.4", "scipy", ]</code></pre> <h3>A) Dataset reading</h3> <div> <div> <pre><code>### Imports import h5py import numpy as np ### Load the input parameters list into a numpy array (shape = (200, 2)) inputf = h5py.File('PPMLES/input_parameters.h5', 'r') input_parameters = np.array((inputf['alpha_inlet'], inputf['friction_velocity'])).T<br>### Load the domain mesh node coordinates<br>meshf = h5py.File('../PPMLES/mesh.h5', 'r')<br>mesh_nodes = np.array((meshf['Nodes']['x'], meshf['Nodes']['y'], meshf['Nodes']['z'])).T&nbsp; ### Load the set of time-averaged LES fields and their associated uncertainty var = 'c' # Can be: 'c', 'u', 'v', 'w', 'crms', 'tke', 'uprim_cprim', 'vprim_cprim', or 'wprim_cprim' fieldsf = h5py.File('PPMLES/ave_fields.h5', 'r') fields_list = fieldsf[var] uncertaintyf = h5py.File('PPMLES/uncertainty_ave_fields.h5', 'r') uncertainty_list = uncertaintyf[var] ### Time series reading example timeseriesf = h5py.File('PPMLES/time_series.h5', 'r') var = 'c' # Can be: 'c', 'u', 'v', or 'w' probe = 32 # Integer between 0 and 92, see probe_network.csv time_list = [] time_series_list = [] for i in range(200): time_list.append(np.array(timeseriesf[f'Sample_{i+1:03}']['time'])) time_series_list.append(np.array(timeseriesf[f'Sample_{i+1:03}'][var][probe]))</code></pre> </div> </div> <h3>B) Interpolation of one-field from the unstructured grid to a new structured grid</h3> <pre><code>### Imports import h5py import numpy as np from scipy.interpolate import griddata ### Load the mean concentration field sample #028 fieldsf = h5py.File('PPMLES/ave_fields.h5', 'r') c = fieldsf['c'][27] ### Load the unstructured grid meshf = h5py.File('PPMLES/mesh.h5', 'r') unstructured_nodes = np.array((meshf['Nodes']['x'], meshf['Nodes']['y'], meshf['Nodes']['z'])).T ### Structured grid definition x0, y0, z0 = -16.9, -115.7, 0. lx, ly, lz = 205.5, 232.1, 20. resolution = 0.75 x_grid, y_grid, z_grid = np.meshgrid(np.linspace(x0, x0 + lx, int(lx/resolution)), np.linspace(y0, y0 + ly, int(ly/resolution)), np.linspace(z0, z0 + lz, int(lz/resolution)), indexing='ij') ### Interpolation of the field on the new grid c_interpolated = griddata(unstructured_nodes, c, (x_grid.flatten(), y_grid.flatten(), z_grid.flatten()), method='nearest')</code></pre> <h3>C) Expression of all time series over the same time window with the same time discretization</h3> <pre><code>### Imports import h5py import numpy as np from scipy.interpolate import griddata ### Define a common time discretization over the 200-s analysis period common_time = np.arange(0., 200., 0.05) u_series_list = np.zeros((200, np.shape(common_time)[0])) ### Interpolate the u-compnent velocity time series at probe DPID10 over this time discretization timeseriesf = h5py.File('PPMLES/time_series.h5', 'r') for i in range(200): sample_time = np.array(timeseriesf[f'Sample_{i+1:03}']['time']) - \ np.array(timeseriesf[f'Sample_{i+1:03}']['Parameters']['t_spinup']) # Offset the spinup time u_series_list[i] = griddata(sample_time, timeseriesf[f'Sample_{i+1:03}']['u'][9], common_time, method='linear')</code></pre> <h3>D) Surrogate model construction example</h3> <p>The training and validation of a POD-GPR surrogate model [Marrel et al. 2015] learning from the PPMLES dataset is given in the following&nbsp;<a href="https://github.com/eliott-lumet/pod_gpr_ppmles">GitHub repository</a>. This surrogate model was successfully used by Lumet et al. 2025 to emulate the LES mean concentration prediction for varying meteorological forcing parameters.</p> <h2>Acknowledgments</h2> <p>This work was granted access to the HPC resources from GENCI-TGCC/CINES (A0062A10822, project 2020-2022). The authors would like to thank Olivier Vermorel for the preliminary development of the LES model, and Simon Lacroix for his proofreading.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types probabilities (part 1)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types probabilities (part 2)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types classification and relative entropy</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Ensemble projections (+ uncertainties) of contemporary (2012-2031) and future (2081-2100) mean annual plankton/phytoplankton/zooplankton species diversity (and species turn-over in time) for the global surface open ocean.

<p><em><strong>Gridded spatial fields (raster objects) containing the species distribution models (SDMs) projections of mean annual plankton total plankton, phytoplankton and zooplankton species diversity from Benedetti et al. (2021). </strong></em></p> <p>The present .grd file (&#39;rasterStack&#39; object in R) contain the fields of mean annual surface plankton/phytoplankton/zooplankton species diversity for the contemporary (2012-2031) and future (2081-2100) conditions of the global open ocean (i.e., data underlying those maps in Figure 1 and Figure 3 of Benedetti et al., 2021). Layers quantifying the uncertainty (i.e., the variablity across models projections estimated through the standard deviation) in ensemble projections were also added (i.e., data underlying the maps in Supplementary Figure 4). See the Methods section of Benedetti et al. (2021) for a full description of the methodology and the ensemble SDMs forecasting framework. The raster layers follow the 1&deg;x1&deg; cell grid of the World Ocean Atlas (https://www.ncei.noaa.gov/).</p> <p>In short, we empirically modelled the monthly and mean annual diversity patterns stemming from the distribution of 860 plankton species (336 phytoplankton, 524 zooplankton) spanning 13 phyla, 71 orders and 324 genera through an ensemble approach based on SDMs. The considered species cover a wide range of traits and functions, representing 10 major plankton functional groups (PFGs; three phytoplankton and seven zooplankton groups). We compiled the species occurrence records from various data sources (available here: https://zenodo.org/record/5101349#.YO7Dqm469lM) and aggregated them onto a monthly-resolved 1&deg;x1&deg; grid, excluding observations from regions where the seafloor is shallower than 200 m. We matched these binned open ocean records with observation-based climatologies of environmental predictors (temperature, dissolved oxygen concentration, solar irradiance, macronutrients concentration, chlorophyll a concentration) that reflect the climatic and biogeochemical conditions of the surface open ocean. Four types of SDMs (generalized linear models, generalized additive models, artificial neural networks, and random forests) were fitted to model the species&rsquo; current environmental habitat suitability patterns. For each SDMs, we used four alternative pools of predictors. Assuming niche conservatism, we projected each of the 16 resulting species-level habitat suitability models into the future using outputs from five ESMs belonging to the Coupled Model Intercomparison Project 5 (CMIP5) that were forced by the Representative Concentration Pathway 8.5 (RCP8.5) scenario of high greenhouse gas concentrations. To this end, we first computed the modelled monthly climatologies of the selected predictors for the 2012-2031 and 2081-2100 periods, and derive the future monthly anomalies from the differences between these two time periods. These anomalies were added to the observation-based monthly climatologies (i.e., those used to train the SDMs) to estimate the future environmental conditions of the ocean, and projected the SDMs in these future conditions. Finally, we estimated the mean annual present and future alpha diversity (species richness; SR) and beta diversity (species turnover through time) patterns for both trophic levels, for each cell, from the ensemble of SDMs. SR ensembles are estimated as the sum of all species&rsquo; habitat suitability patterns averaged across all 80 possible combinations (i.e., &quot;ensemble members&quot;) of SDMs (n = 4), ESMs (n = 5) and predictor pools (n = 4). To assess the uncertainties of our diversity projections based on the ensemble members, we compute the interquartile range of the 80 ensemble members SR projections. We calculate species turnover as the change in mean annual species composition between present and future time based on Jaccard&rsquo;s dissimilarity index and by decomposing this total turnover into the true species turnover (ST, also known as species replacement) and the nestedness (SR change) components. Numerous tests are conducted to ensure the robustness of the results with regard to the spatially and temporally highly uneven sampling effort as well as with regard to the relative role of different predictors.</p> <p><strong>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 862923. This output reflects only the author&rsquo;s view, and the European Union cannot be held responsible for any use that may be made of the information contained therein.</strong></p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CRCM5-CMIP6 : A dynamically-downscaled ensemble of CMIP6 simulations.

<h1>CRCM-CMIP</h1> <h2>Data reference</h2> <p>Paquin, D., C. McCray, C. B. Gauthier, M. Gigu&egrave;re, O. Asselin, P .Bourgault, M.-P. Labont&eacute; and D. Matte. The CRCM5-CMIP6 Ouranos&rsquo; ensemble : A dynamically-downscaled ensemble of CMIP6 simulations over North America. Accepted in Scientific Data.</p> <p><a href="https://www.ouranos.ca/en">Ouranos</a> : Canadian Regional Climate Model &ndash; version 5</p> <p><strong>Martynov et al. 2013, Separovic et al. 2013</strong></p> <p>Based on GEM 3.3.3.1</p> <h3>Configuration</h3> <p>NAM-11 CORDEX North American domain at 0.11&deg; 695x668 grid points including a 20-point sponge (and halo) zone surrounding the domain, 5-minute time steps, xlat1=28.525 xlon2=145.955. 56 vertical levels and a top at 10 hPa. 17 surface levels and a bottom at 15 m.</p> <h3>Spectral Nudging</h3> <p>A spectral nudging is applied to the horizontal wind component with a half-response wavelength of 1177km and a relaxation time of 13.34 h. The nudging strength is set to zero from the surface to a height of 500 hPa and increases linearly onward to the top of the model&rsquo;s simulated atmosphere (10 hPa).</p> <h2>Parameterization</h2> <h3>Atmosphere</h3> <p>Precipitation: modified Sundqvist &nbsp;(1998); precipitation partition Bourgouin &nbsp;(2000) ; Implicit vertical diffusion.&nbsp;<br>Shallow convection: Kuo (1965) transient shallow, Non‐cloudy boundary layer formulation.&nbsp;<br>Deep convection: Kain-Fritsch (1990);&nbsp;<br>Radiation: Li &amp; Barker (2005)</p> <h3>Surface</h3> <p>CLASS3.5c (Verseghy, 1993)</p> <p>Lake model: FLake</p> <h3>Ocean</h3> <p>Prescribed SST &amp; sea ice fraction</p> <h3>Aerosol</h3> <p>Prescribed</p> <h2>Data Access</h2> <p>Due to its large size, the full dataset can't yet be shared publicly.</p> <p>A subset of the variables are stored on Ouranos' THREDDS server.</p> <p>- Annual files : <a href="https://pavics.ouranos.ca/twitcher/ows/proxy/thredds/catalog/birdhouse/disk2/ouranos/CORDEX/catalog.html">https://pavics.ouranos.ca/twitcher/ows/proxy/thredds/catalog/birdhouse/disk2/ouranos/CORDEX/catalog.html</a><br>- Aggregated datasets : <a href="https://pavics.ouranos.ca/twitcher/ows/proxy/thredds/catalog/datasets/simulations/RCM-CMIP6/catalog.html">https://pavics.ouranos.ca/twitcher/ows/proxy/thredds/catalog/datasets/simulations/RCM-CMIP6/catalog.html</a></p> <p>Other variables can be provided upon request by writing to simulations_ouranos@ouranos.ca.</p> <p>All data are available through a&nbsp;<a href="https://creativecommons.org/licenses/by/4.0/legalcode">CC-BY 4.0</a> license.</p> <h2>Acknowlegments</h2> <p>Developed by the&nbsp;<a href="https://escer.uqam.ca/">ESCER Centre</a> at UQAM (Universit&eacute; du Qu&eacute;bec &agrave; Montr&eacute;al) with the collaboration of Environment and Climate Change Canada (ECCC).&nbsp;<strong>CRCM5; Martynov et al. 2013, Separovic et al. 2013</strong></p> <p>The CRCM5 data has been generated and supplied by Ouranos.</p> <p>CRCM5 computations were made on the supercomputers beluga and narval managed by Calcul Qu&eacute;bec and the&nbsp;<a href="https://alliancecan.ca/en">Digital Research Alliance of Canada</a>. The operation of this supercomputer received financial support from Innovation, Science and Economic Development Canada and the Minist&egrave;re de l&rsquo;&Eacute;conomie et de l&rsquo;Innovation du Qu&eacute;bec.</p> <h2>Some references for CRCM5</h2> <p>Asselin, M. Leduc, D. Paquin, K. Winger, A. Di Luca, M. Bukovsky, B. Music, and M. Gigu&egrave;re (2022). On the Intercontinental Transferability of Regional Climate Model Response to Severe Forestation. &nbsp;MDPI's Climate&nbsp;<br><a href="https://doi.org/10.3390/cli10100138">https://doi.org/10.3390/cli10100138</a>&nbsp;</p> <p>Bresson, E., R. Laprise, D. Paquin, J. M. Th&eacute;riault, R. de Elia, 2017: Evaluating CRCM5 ability to simulate mixed precipitation. Atmosphere-Ocean. 55(2); 79-93.&nbsp;<a href="http://dx.doi.org/10.1080/07055900.2017.1310084">http://dx.doi.org/10.1080/07055900.2017.1310084</a>&nbsp;</p> <p>Leduc, M., A. Mailhot, A. Frigon, J.-L. Martel, R. Ludwig, G.B. Brietzke, M. Gigu&egrave;re, F. Brissette, R. Turcotte, M. Braun, (2019) ClimEx project: a 50-member ensemble of climate change projections at 12-km resolution over Europe and northeastern North America with the Canadian Regional Climate Model (CRCM5). Journal of Applied Meteorology and Climatology.&nbsp;<a href="https://doi.org/10.1175/JAMC-D-18-0021.1" target="_blank" rel="noopener">https://doi.org/10.1175/JAMC-D-18-0021.1</a></p> <p>Martynov A, R Laprise, L Sushama, K Winger, L Separovic, B Dugas. 2013. Reanalysis-driven climate simulation over CORDEX North America domain using the Canadian Regional Climate Model, version 5: model performance evaluation. Clim Dyn 41:2973-3005.&nbsp;<a href="https://doi.org/10.1007/s00382-013-1778-9">https://doi.org/10.1007/s00382-013-1778-9</a></p> <p>Martynov A, L Sushama, R Laprise, K Winger, B Dugas. 2012. Interactive lakes in the Canadian regional climate model version 5: the role of lakes in the regional climate of North America. Tellus A 64, 016226. <a href="https://doi.org/10.3402/tellusa.v64i0.16226">https://doi.org/10.3402/tellusa.v64i0.16226</a>.</p> <p>Martynov A, L Sushama, R Laprise. 2010. Simulation of temperate freezing lakes by one-dimensional lake models: performance assessment for interactive coupling with regional climate models. Boreal Env Res 15:143-164.</p> <p>Matte, D., Th&eacute;riault, J. M., &amp; Laprise, R. (2019). Mixed precipitation occurrences over southern Qu&eacute;bec, Canada, under warmer climate conditions using a regional climate model. Climate Dynamics, 53(1), 1125&ndash;1141. <a href="https://doi.org/10.1007/s00382-018-4231-2">https://doi.org/10.1007/s00382-018-4231-2</a></p> <p>McCray, C. D., D. Paquin, J. M. Th&eacute;riault, &Eacute;. Bresson (2022). A multi-algorithm analysis of projected changes to freezing rain over North America in an ensemble of regional climate model simulations. Journal of Geophysical Research -Atmospheres <a href="https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2022JD036935">https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2022JD036935</a></p> <p>McCray, D. C., J. M. Th&eacute;riault, D. Paquin, &Eacute;. Bresson, 2022. Quantifying the impact of precipitation-type algorithm selection on the representation of freezing rain in an ensemble of regional climate model simulations. Journal of Applied Meteorology and Climatology. <a href="https://journals.ametsoc.org/view/journals/apme/aop/JAMC-D-21-0202.1/JAMC-D-21-0202.1.xml">https://journals.ametsoc.org/view/journals/apme/aop/JAMC-D-21-0202.1/JAMC-D-21-0202.1.xml</a>&nbsp;</p> <p>McCray, C.D., G. Schmidt, D. Paquin, M. Leduc, Z. Bi, M. Radiyat, C. Silverman, M. Spitz, B. Brettschneider (2023). Changing Nature of High-Impact Snowfall Events in Eastern North America. Journal of Geophysical Research: Atmospheres. <a href="https://doi.org/10.1029/2023JD038804">https://doi.org/10.1029/2023JD038804</a></p> <p>Mironov D, E Heise, E Kourzeneva, B Ritter, N Schneider, A Terzhevik. 2010. Implementation of the lake parameterisation scheme FLake into the numerical weather prediction model COSMO. Boreal Env Res 15:218-230.</p> <p>Mittermeier, M., E. Bresson, D. Paquin, R. Ludwig, 2021 A deep learning approach for the identification of long-duration mixed precipitation in Montr&eacute;al (Canada). Atmosphere-Ocean. <a href="https://doi.org/10.1080/07055900.2021.1992341">https://doi.org/10.1080/07055900.2021.1992341</a></p> <p>Riette S, D Caya. 2002. Sensitivity of short simulations to the various parameters in the new CRCM spectral nudging. &ndash; In: RITCHIE, H. (Ed.): Research activities in Atmospheric and Oceanic Modeling, WMO/TD No. 1105, Report No. 32: 7.39&ndash;7.40.</p> <p>P&eacute;rez Bello, A., A. Mailhot and D. Paquin, 2021 The response of daily and sub-daily extreme precipitations to changes in surface and dew point temperatures. Journal of Geophysical Research &ndash; Atmospheres <a href="http://dx.doi.org/10.1029/2021JD034972">http://dx.doi.org/10.1029/2021JD034972</a></p> <p>P&eacute;rez Bello, A., A. Mailhot, D. Paquin and D. Paquin-Ricard (2022). Temperature-precipitation scaling rates: a rainfall event-based perspective. Journal of Geophysical Research &ndash; Atmospheres. <a href="https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2022JD037873">https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2022JD037873</a></p> <p>Separovic L, A Alexandru, R Laprise, A Martynov, L Sushama, K Winger, K Tete, M Valin. 2013. Present climate and climate change over North America as simulated by the fifth-generation Canadian regional climate model. Clim Dyn 41:3167-3201. <a href="https://doi.org/10.1007/s00382-013-1737-5">DOI 10.1007/s00382-013-1737-5</a>.</p> <p>St-Pierre, M., J. Th&eacute;riault and D. Paquin, 2019. Influence of the model spatial resolution on atmospheric conditions leading to freezing rain in regional climate simulations. Atmosphere-Ocean, <a href="https://doi.org/10.1007/s00382-013-1737-5">https://doi.org/10.1080/07055900.2019.1583088</a>.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

The convergence of an ensemble to an attractor during climate change

<p>Time series of the annual mean near-surface air temperature of a grid point in the Southern Pacific in 192 members in each of two ensembles simulated by the Planet Simulator (Fraedrich et al., 2005) as described in Drotos et al. (2017) in relation to Fig. 2. One ensemble was initialised in year 0 (but note that the time series are available only from year 500 on) and the other in year 610 (the temperature of the last unperturbed year, year 609, is also included in the first lines of the corresponding files). The latter date is during a simulated climate change.</p> <p>G. Dr&oacute;tos, T. B&oacute;dai and T. T&eacute;l (2017), &quot;On the importance of the convergence to climate attractors&quot;. Eur. Phys. J. Spec. Top. 226, 2031&ndash;2038. https://doi.org/10.1140/epjst/e2017-70045-7<br> K. Fraedrich, H. Jansen, E. Kirk, U. Luksch, F. Lunkeit (2005), &quot;The Planet Simulator: Towards a user friendly model&quot;. Meteorol. Z. 14, 299&ndash;304. https://doi.org/10.1127/0941-2948/2005/0043</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Database of small molecule X-ray absorption spectra, featurized structures, and neural network ensembles

<p>Companion data for arXiv preprint <em>Uncertainty-aware predictions of molecular X-ray absorption spectra using neural network ensembles</em>&nbsp;(<a href="https://arxiv.org/abs/2210.00336">https://arxiv.org/abs/2210.00336</a>), by&nbsp;Animesh Ghose, Mikhail Segal, Fanchen Meng, Zhu Liang, Mark S. Hybertsen, Xiaohui Qu, Eli Stavitski, Shinjae Yoo, Deyu Lu &amp;&nbsp;Matthew R. Carbone.</p> <p><strong>Included</strong></p> <ul> <li>*-XANES-*.tar.bz2: raw&nbsp;input/output files for all molecular simulations used in the work. These inputs and outputs correspond to the structural data in the QM9 dataset.</li> <li>ml_ready.tar.bz2: machine learning-ready data (featurized spectra). Used as input to the neural network ensembles.</li> <li>XANES-220712-ACSF-*.tar.bz2: neural network ensembles used in this work.</li> </ul> <p><strong>Notes</strong></p> <ul> <li>The FEFF9 code [J. J. Rehr, J. J. Kas, F. D. Vila, M. P. Prange, and&nbsp;K. Jorissen, <em>Phys. Chem. Chem. Phys.</em> <strong>12</strong>, 5503 (2010)]&nbsp;was used to generate all X-ray absorption near-edge structure (XANES) spectra.</li> <li>All molecular structures were sourced from the QM9 database [R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. Von Lilienfeld, <em>Sci. Data</em> <strong>1</strong>, 1 (2014)].</li> </ul> <p><strong>Funding</strong></p> <p>This research is based upon work supported by the U.S. Department of Energy, Office of Science, Office Basic Energy Sciences, under Award Number FWP PS-030. This research also used theory and computational resources of the Center for Functional Nanomaterials, which is a U.S. Department of Energy Office of Science User Facility, and the Scientific Data and Computing Center, a component of the Computational Science Initiative, at Brookhaven National Laboratory under Contract No. DE-SC0012704.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

SUMMA/mizuRoute model configurations, parameters, and ensemble statistics for representative cryosphere basins

<p>Meteorological forcing is a major source of uncertainty in hydrological modeling. The recent development of probabilistic large-domain meteorological datasets enables convenient uncertainty characterization, which however is rarely explored in large-domain research.&nbsp;Tang et al. (2023)&nbsp;analyze&nbsp;how uncertainties in meteorological forcing data affect hydrological modeling in 289 representative cryosphere basins by forcing the Structure for Unifying Multiple Modeling Alternatives (SUMMA) and mizuRoute models with precipitation and air temperature ensembles from the Ensemble Meteorological Dataset for Planet Earth (EM-Earth).&nbsp;EM-Earth probabilistic estimates are used in ensemble simulation for uncertainty analysis. The results reveal the magnitude, spatial distribution, and scale effect of uncertainties in meteorological, snow, runoff, soil water, and energy variables.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Cloud Botany LES ensemble visualizations

<p><strong>Animations and graphs of the Cloud Botany ensemble of large eddy simulations</strong></p> <p>The data is stored as a single zip file. When unpacked, the visualizations can be navigated in a web browser - start by opening index.html in the top directory.</p> <p>The simulations were performed with DALES, the Dutch Atmospheric Large Eddy Simulation, on Supercomputer Fugaku. For access to the data itself, see <a href="https://howto.eurec4a.eu/botany_dales.html">How To EUREC4A.</a></p> <p><strong>Additional material</strong>: ERA5 data for the experiment region for the spring 2020 and JOANNE dropsonde data from the EUREC4A campaign, used to define the ranges for the Cloud Botany ensemble parameter space.</p>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record