Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

30

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

30 results for “large ensemble”

Learn how ShareScore rates datasets ↗
zenodo52/100

A large ensemble of CMIP6-based transient climate scenarios for impact assessment in Great Britain.

<p>Climate change impact assessments often require a large ensemble of local-scale transient climate scenarios. Each ensemble member represents plausible long weather series at a local scale. The climate projections from Global Climate Models (GCMs) are difficult to use at local scale due to their coarse spatial and temporal resolution. Moreover, very few projections are usually available for each GCM due to a high computational cost. An alternative approach involves employing a stochastic weather generator to produce a large number of transient scenarios based on the climate projections from GCMs. In a current dataset, transient climate scenarios were generated using the LARS-WG weather generator, based on climate projections from &nbsp;GCMs from the CMIP6 ensemble across 26 representative sites throughout the UK. Each transient scenario spans the period from 2020 to 2090.&nbsp; At each site, 100 transient scenarios were generated for two emission scenarios (SSP2-4.5 and SSP5-8.5) and five selected GCMs from CMIP6 (ACCESS-ESM1-5, CNRM-CM6-1, HadGEM3-GC31-LL, MPI-ESM1-2-LR, and MRI-ESM2-0). The choice of GCMs were&nbsp; based on their performance over northern Europe and their climate sensitivity. The use of a subset of GCMs substantially reduces computational time required for impact assessment, while allowing to quantify uncertainties in impacts related to uncertain future climate. The dataset can be used with impact models in various fields, including, land and water resources, agriculture and food production, ecology and epidemiology, and human health and welfare, when undertaking impact assessment of climate change and decision support for mitigation and adaptation.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Dataset for "Remapping of Greenland ice sheet surface mass balance anomalies for large ensemble sea-level change projections"

<p>This dataset is used to reproduce the results presented in the following publication:</p> <p>Goelzer, H., Noel, B. P. Y., Edwards, T. L., Fettweis, X., Gregory, J. M., Lipscomb, W. H., van de Wal, R. S. W., and van den Broeke, M. R.: Remapping of Greenland ice sheet surface mass balance anomalies for large ensemble sea-level change projections, The Cryosphere Discuss., https://doi.org/10.5194/tc-2019-188, in review, 2019.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo48/100

BioVars - bioclimatic datasets for Europe based on a large regional climate ensemble for periods between 1971 to 2098

<p>We present 26 bio-climatic variables that are calculated based on a large ensemble consisting of 70 bias-adjusted GCM-RCM (Global Climate Model &ndash; Regional Climate Model) simulations for 1971 to 2098. Both, the historic and the projection periods were calculated using the same models to ensure consistency between the periods. The variables are validated against E-OBS observations from which we calculated the same bio-climatic variables. For projection periods we chose 20 year ranges between 2021 to 2098. Here, we offer two versions of them 1) variables separated into RCP 2.6, 4.5 and 8.5 including the 5th, 50th and 95th percentiles among the realisations and within the RCPS. And 2) variables per realisation separately. We then extracted the temporal 5th, 50th and 95th percentile per period as representing values. Each zipped file contains these 26 bio-climatic variables according to their aggregation. The variables and the units are explained within the data descriptor publication.&nbsp;</p> <p>&nbsp;</p> <p><strong>File descriptions</strong></p> <ul> <li>bioVars_1971-2000_met.tar.gz &gt;&gt; Projections per realisations for period 1971-2000</li> <li>bioVars_2021-2040_met.tar.gz &gt;&gt; Projections per realisations for period 2021-2040</li> <li>bioVars_2041-2060_met.tar.gz &gt;&gt; Projections per realisations for period 2041-2060</li> <li>bioVars_2061-2080_met.tar.gz &gt;&gt; Projections per realisations for period 2061-2080</li> <li>bioVars_2079-2098_met.tar.gz &gt;&gt; Projections per realisations for period 2079-2098</li> <li>bioVars_2021-2040_rcp.tar.gz &gt;&gt; Projections per RCP for period 2021-2040</li> <li>bioVars_2041-2060_rcp.tar.gz &gt;&gt; Projections per RCP for period 2041-2060</li> <li>bioVars_2061-2080_rcp.tar.gz &gt;&gt; Projections per RCP for period 2061-2080</li> <li>bioVars_2079-2098_rcp.tar.gz &gt;&gt; Projections per RCP for period 2079-2098</li> <li>validation.tar.gz &gt;&gt; Validation using E-OBS (v20.0) and Worldclim (version 2.1)</li> </ul> <p>&nbsp;</p> <p><strong>References</strong> <br>Reichmuth, A., Rakovec, O., Boeing, F. <em>et al.</em> BioVars - A bioclimatic dataset for Europe based on a large regional climate ensemble for periods in 1971&ndash;2098. <em>Sci Data</em> <strong>12</strong>, 217 (2025). https://doi.org/10.1038/s41597-025-04507-w</p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Graph Data: Hydrological impact of widespread afforestation in Great Britain using a large ensemble of modelled scenarios

<p>Data used for creating the figures in the paper:&nbsp;Hydrological impact of widespread afforestation in Great Britain using a large ensemble of modelled scenarios.</p> <p>It contains the&nbsp;flow exceedances (as mm day<sup>-1</sup>),&nbsp;flow duration slope, median elasticity&nbsp;and runoff ratio for the different afforestation scenarios. Also included is the information on the changes of broadleaf afforestation.&nbsp;</p> <p>If you have any questions, please email marcus.buechel@ouce.ox.ac.uk.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Large Ensemble Dataset for Discovering Global Peak Water Limit of Future Groundwater Withdrawals Using 900 GCAM Runs

<h2><strong>Global Groundwater Withdrawals Peak&nbsp;Over the 21st Century&nbsp;</strong></h2> <p>The large ensemble dataset contains groundwater related model outputs from 900 scenarios modeled using <a href="http://jgcri.github.io/gcam-doc/toc.html">Global Change Analysis Model (GCAM)</a>. The scenario ensemble&nbsp;members include five Shared Socioeconomic Pathways (SSPs), four Representative Concentration Pathways (RCPs), five global climate model outputs, three groundwater depletion limits, two surface water storage expansion regimes, and two historical groundwater depletion trends.</p> <h3><strong>Journal Article</strong></h3> <p>Niazi, H., Wild, T.B., Turner, S.W.D., Graham, N.T., Hejazi, M., Msangi, S., Kim, S., Lamontagne, J.R., &amp; Zhao, M. (2024).&nbsp;<a href="https://rdcu.be/dFpb5">Global peak water limit of future groundwater withdrawals</a>.&nbsp;<em>Nature Sustainability, 7</em>(4), 413&ndash;422.&nbsp;<a href="https://doi.org/10.1038/s41893-024-01306-w" rel="nofollow">https://doi.org/10.1038/s41893-024-01306-w</a></p> <p>Read full-text here: <a href="https://rdcu.be/dFpb5">https://rdcu.be/dFpb5</a>&nbsp;</p> <h3><strong>Data Repository&nbsp;</strong></h3> <p>This <em><strong>data</strong></em> repository is to be used in combination with the&nbsp;<em><strong>main</strong></em>&nbsp;<a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">meta-repository</a> containing all scripts and files for reproducing the experiment as well as the analysis and post-processing of the model outputs.</p> <p>Scripts and smaller files are provided in the <a href="https://github.com/JGCRI/niazi-etal_202X_xyz">GitHub meta-repository</a> whereas larger files are provided in this data repository. Please complete the repository by placing the files as described hereunder. Please find the GitHub meta-repository here: <a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">https://github.com/JGCRI/niazi-etal_2024_nature-sustainability</a></p> <p>Descriptions of files:</p> <ol> <li><em><strong>gcam-5.7z</strong></em> contains the GCAM version used to simulate&nbsp;900 scenarios of plausible futures. The model folder contains all necessary input files to reproduce the simulations. <ul> <li>The model is to be used in combination with the <a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">meta-repository</a>&nbsp;to setup batch runs on cluster.</li> <li>Please navigate to <a href="https://github.com/JGCRI/niazi-etal_202X_xyz/tree/main/model">model/</a> folder for&nbsp;other scenario-specific and model setup folders and files. <em><strong>gcam-5</strong></em>&nbsp;is to be extracted in the same directory (./<em>model/gcam-5/</em>).&nbsp;</li> <li>For the first-time users of GCAM, please follow&nbsp;guidance on <a href="http://jgcri.github.io/gcam-doc/toc.html">GCAM wiki</a>&nbsp;to setup GCAM or for background knowledge.&nbsp;</li> </ul> </li> <li><em><strong>crop_yeild.7z</strong></em>: This file contains inputs related to&nbsp;climate impacts on crop yields. This is to be downloaded and extracted&nbsp;in the&nbsp;<a href="https://github.com/JGCRI/niazi-etal_202X_xyz/tree/main/model/combined_impacts">model/combined_impacts/</a>&nbsp;folder.&nbsp;</li> <li><em><strong>outputs-all.7z: </strong></em>Key model outputs queried and collated from 900 GCAM runs are explained hereunder.&nbsp;The files could be downloaded individually (.csv&nbsp;files)&nbsp;or all at once in .7z format (<a href="../api/files/80b237d3-b22f-499f-8b8e-76c3846720a0/outputs-all.7z">outputs-all.7z</a>). These files are to be placed in the <a href="https://github.com/JGCRI/niazi-etal_202X_xyz/tree/main/model/outputs">model/outputs</a>&nbsp;folder of the <a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">meta-repository</a>.&nbsp; <ul> <li><em><strong>ag_prod_all_GW_scenarios.csv</strong></em>&nbsp;- Agricultural production across all scenario for 2050 and 2100 (tonnes)</li> <li><em><strong>prices_water_withdrawal_all.csv</strong> -&nbsp;</em>Water prices across all scenarios and years ($/km<sup>3</sup>)</li> <li><em><strong>global_irrigated_prod_by_crop.csv</strong></em>&nbsp;-&nbsp;All irrigated agricultural production for each crop across and scenarios all years (tonnes)</li> <li><em><strong>surface_water_production_all.csv</strong></em>&nbsp;- Runoff across all scenarios and years (km<sup>3</sup>)</li> <li><em><strong>groundwater_production_FINAL.csv</strong></em>&nbsp;- Groundwater withdrawals across all scenarios and years (km<sup>3</sup>)</li> <li><em><strong>water_withdrawals_desal_all.csv</strong></em>&nbsp;- Water withdrawals from desalination plants across all scenarios and years (km<sup>3</sup>)</li> </ul> </li> </ol> <h3><strong>Short introduction to the study</strong></h3> <p>Using 900 GCAM runs, this study finds that global groundwater withdrawals are expected to peak around mid-century, followed by a decline through 21st century, exposing about half of the population living in one-third of basins to groundwater stress, with cost and availability of surface water storage being the most significant driver of future groundwater withdrawals. This first-ever robust, quantitative confirmation of the peak-and-decline pattern for groundwater, previously only known for fossil fuels and minerals, raises concerns for basins heavily dependent on groundwater.</p> <p>Niazi, H., Wild, T.B., Turner, S.W.D., Graham, N.T., Hejazi, M., Msangi, S., Kim, S., Lamontagne, J.R., &amp; Zhao, M. (2024).&nbsp;<a href="https://rdcu.be/dFpb5">Global peak water limit of future groundwater withdrawals</a>.&nbsp;<em>Nature Sustainability, 7</em>(4), 413&ndash;422.&nbsp;<a href="https://doi.org/10.1038/s41893-024-01306-w" rel="nofollow">https://doi.org/10.1038/s41893-024-01306-w</a></p> <p>Read full-text here: <a href="https://rdcu.be/dFpb5">https://rdcu.be/dFpb5</a></p> <h3><strong>Contact&nbsp;</strong></h3> <p>Please reach out to Hassan Niazi at&nbsp;<a href="mailto:hassan.niazi@pnnl.gov">hassan.niazi@pnnl.gov</a> for any questions.&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

PPMLES – Perturbed-Parameter ensemble of MUST Large-Eddy Simulations

<h2>Dataset description</h2> <p>This repository contains the PPMLES (Perturbed-Parameter ensemble of MUST Large-Eddy Simulations) dataset, which corresponds to the main outputs of 200 large-eddy simulations (LES) of microscale pollutant dispersion that replicate the MUST field experiment [Biltoft. 2001, Yee and Biltoft. 2004] for varying meteorological forcing parameters.</p> <p>The goal of the PPMLES dataset is to provide a comprehensive dataset to better understand the complex interactions between the atmospheric boundary layer (ABL), the urban environment, and pollutant dispersion. It was originally used to assess the impact of the meteorological uncertainty on microscale pollutant prediction and to build a surrogate model that can replace the costly LES model [Lumet et al. 2025]. The total computational cost of the PPMLES dataset is estimated to be about 6 million core hours.</p> <p>For each sample of meteorological forcing parameters (inlet wind direction and friction velocity), the <a href="https://www.cerfacs.fr/avbp7x/">AVBP</a> solver code [Schonfeld and Rudgyard. 1999, Gicquel et al. 2011] was used to perform LES at very high spatio-temporal resolution (1e-3s time step, 30cm discretization length) to provide a fine representation of the pollutant concentration and wind velocity statistics within the urban-like canopy. The total computational cost of the PPMLES dataset is estimated to be about 6 million core hours.</p> <h2>File list</h2> <p>The data is stored in <a href="https://www.hdfgroup.org/solutions/hdf5/">HDF5</a> files, which can be efficiently processed in Python using the <a href="https://docs.h5py.org/en/stable/index.html">h5py</a> module.&nbsp;</p> <ul> <li><em>input_parameters.h5:</em> list of the 200 input parameter&nbsp;samples<em>&nbsp;(alpha_inlet, ustar) </em>obtained using the&nbsp;Halton sequence that defines the PPMLES ensemble.</li> <li><em>ave_fields.h5</em>: lists of the main field statistics predicted by each of the 200 LES samples over the 200-s reference window [Yee and Biltoft. 2004], including: <ul> <li><em>c:</em> the time-averaged pollutant concentration in ppmv <em>(dim = (n_samples, n_nodes) = (200, 1878585))</em>,&nbsp;</li> <li><em>(u, v, w):&nbsp;</em>the time-averaged wind velocity components in m/s,</li> <li><em>crms: </em>the root mean square concentration fluctuations in ppmv,&nbsp;</li> <li><em>tke:</em> the turbulent kinetic energy in m^2/s^2,</li> <li><em>(uprim_cprim, vprim_cprim, wprim_cprim)</em>: the pollutant turbulent transport components</li> </ul> </li> <li><em>uncertainty.h5</em>: lists of&nbsp;the estimated aleatory uncertainty induced by the internal variability of the LES&nbsp;<em>(variability_#)</em> [Lumet et al. 2024] for each of the fields in <em>ave_fields.h5</em>. Also includes the stationary bootstrap [Politis and Romano. 1994] parameters <em>(n_replicates, block_length) </em>used to estimate the uncertainty for each field and each sample.</li> <li><em>mesh.h5</em>: the tetrahedral mesh on which the fields are discretized, composed of about 1.8 millions of nodes.</li> <li><em>time_series.h5</em>:&nbsp;HDF5 file consisting of 200 groups (<em>Sample_NNN</em>) each containing the time series of the pollutant concentration (c) and wind velocity components (u, v, w) predicted by the LES sample #NNN at 93 locations.&nbsp;</li> <li><em>probe_network.dat</em>: provides the location of each of the 93 probes corresponding to the positions of the experimental campaign sensors [Biltoft. 2001].</li> </ul> <p><strong>Warning:</strong> the propylene concentration are expressed in ppmv, except in time_series.h5 in which they are given as mass fractions. To convert them in ppmv, the formula is: <code>c = c * (rho/rho_propylene) * 10**6</code> with (rho/rho_propylene) = 0.66 the density ratio between air and propylene.</p> <h2>Code examples</h2> <p>In the following, examples of how to use the PPMLES dataset in Python are provided. These examples have the following dependencies:&nbsp;</p> <pre><code>requires-python = "&gt;=3.9" dependencies = [ "h5py==3.8.0", "numpy==1.26.4", "scipy", ]</code></pre> <h3>A) Dataset reading</h3> <div> <div> <pre><code>### Imports import h5py import numpy as np ### Load the input parameters list into a numpy array (shape = (200, 2)) inputf = h5py.File('PPMLES/input_parameters.h5', 'r') input_parameters = np.array((inputf['alpha_inlet'], inputf['friction_velocity'])).T<br>### Load the domain mesh node coordinates<br>meshf = h5py.File('../PPMLES/mesh.h5', 'r')<br>mesh_nodes = np.array((meshf['Nodes']['x'], meshf['Nodes']['y'], meshf['Nodes']['z'])).T&nbsp; ### Load the set of time-averaged LES fields and their associated uncertainty var = 'c' # Can be: 'c', 'u', 'v', 'w', 'crms', 'tke', 'uprim_cprim', 'vprim_cprim', or 'wprim_cprim' fieldsf = h5py.File('PPMLES/ave_fields.h5', 'r') fields_list = fieldsf[var] uncertaintyf = h5py.File('PPMLES/uncertainty_ave_fields.h5', 'r') uncertainty_list = uncertaintyf[var] ### Time series reading example timeseriesf = h5py.File('PPMLES/time_series.h5', 'r') var = 'c' # Can be: 'c', 'u', 'v', or 'w' probe = 32 # Integer between 0 and 92, see probe_network.csv time_list = [] time_series_list = [] for i in range(200): time_list.append(np.array(timeseriesf[f'Sample_{i+1:03}']['time'])) time_series_list.append(np.array(timeseriesf[f'Sample_{i+1:03}'][var][probe]))</code></pre> </div> </div> <h3>B) Interpolation of one-field from the unstructured grid to a new structured grid</h3> <pre><code>### Imports import h5py import numpy as np from scipy.interpolate import griddata ### Load the mean concentration field sample #028 fieldsf = h5py.File('PPMLES/ave_fields.h5', 'r') c = fieldsf['c'][27] ### Load the unstructured grid meshf = h5py.File('PPMLES/mesh.h5', 'r') unstructured_nodes = np.array((meshf['Nodes']['x'], meshf['Nodes']['y'], meshf['Nodes']['z'])).T ### Structured grid definition x0, y0, z0 = -16.9, -115.7, 0. lx, ly, lz = 205.5, 232.1, 20. resolution = 0.75 x_grid, y_grid, z_grid = np.meshgrid(np.linspace(x0, x0 + lx, int(lx/resolution)), np.linspace(y0, y0 + ly, int(ly/resolution)), np.linspace(z0, z0 + lz, int(lz/resolution)), indexing='ij') ### Interpolation of the field on the new grid c_interpolated = griddata(unstructured_nodes, c, (x_grid.flatten(), y_grid.flatten(), z_grid.flatten()), method='nearest')</code></pre> <h3>C) Expression of all time series over the same time window with the same time discretization</h3> <pre><code>### Imports import h5py import numpy as np from scipy.interpolate import griddata ### Define a common time discretization over the 200-s analysis period common_time = np.arange(0., 200., 0.05) u_series_list = np.zeros((200, np.shape(common_time)[0])) ### Interpolate the u-compnent velocity time series at probe DPID10 over this time discretization timeseriesf = h5py.File('PPMLES/time_series.h5', 'r') for i in range(200): sample_time = np.array(timeseriesf[f'Sample_{i+1:03}']['time']) - \ np.array(timeseriesf[f'Sample_{i+1:03}']['Parameters']['t_spinup']) # Offset the spinup time u_series_list[i] = griddata(sample_time, timeseriesf[f'Sample_{i+1:03}']['u'][9], common_time, method='linear')</code></pre> <h3>D) Surrogate model construction example</h3> <p>The training and validation of a POD-GPR surrogate model [Marrel et al. 2015] learning from the PPMLES dataset is given in the following&nbsp;<a href="https://github.com/eliott-lumet/pod_gpr_ppmles">GitHub repository</a>. This surrogate model was successfully used by Lumet et al. 2025 to emulate the LES mean concentration prediction for varying meteorological forcing parameters.</p> <h2>Acknowledgments</h2> <p>This work was granted access to the HPC resources from GENCI-TGCC/CINES (A0062A10822, project 2020-2022). The authors would like to thank Olivier Vermorel for the preliminary development of the LES model, and Simon Lacroix for his proofreading.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

SOCAT+USV sampling masks for ML reconstruction of surface ocean pCO2 using the Large Ensemble Testbed

<p>Here we provide sampling masks used in the study "Assessing improvements in global ocean pCO2 machine learning reconstructions with Southern Ocean autonomous sampling" (Heimdal et al., 2023, https://doi.org/10.5194/bg-2023-160). In this paper, we reconstruct surface ocean pCO2 using the Large Ensemble Testbed (Gloege et al., 2021, https://doi.org/10.1029/2020GB006788) and the pCO2-Residual method (Bennington et al., 2022, https://doi.org/10.1029/2021MS002960). We provide 11 different sampling masks that correspond to the experiments presented in Heimdal et al. (2023), which include different sampling patterns of USV Saildrones in the Southern Ocean (SOCAT+USV sampling).</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Spatial patterns of extreme precipitation and their changes under ~2 °C global warming: A large-ensemble study of the western US: Data Release

<p>This dataset supports the analysis in Rupp et al. (2022). The dataset consists of 17,223 data files containing the water year (WY) maximum of the daily-averaged precipitation rate simulated with the HadRM3p regional climate model configured for the western United States. Each file contains the WY maxima across the model domain for a single WY, single model parameterization, and single set of initial conditions. Please refer to Hawkins et al. (2019) and Rupp et al. (2022) for a description of how the climate model data were generated.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Data used in a manuscript entitled "Large ensemble simulation for investigating predictability of precursor vortices of Typhoon Faxai in 2019 with a 14-km mesh global nonhydrostatic atmospheric model" submitted to Geophysical Research Letters

<p>This include a dataset used in a manuscript entitled &ldquo;Large ensemble simulation for investigating predictability of precursor vortices of Typhoon Faxai in 2019 with a 14-km mesh global nonhydrostatic atmospheric model&rdquo; by Yamada and co-authors, which is submitted to Geophysical Research Letters.</p> <p>Contact: Yohei Yamada (yoheiy@jamstec.go.jp)</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Model output used in the manuscript "Micro and macro parametric uncertainty in climate change prediction: a large ensemble perspective"

<p>This *.zip file contains the model output from ensemble simulations for the Lorenz 84-Stommel 61 model (hereafter L84-S61; <a href="https://doi.org/10.3402/tellusa.v53i5.12229" target="_blank" rel="noopener">Van Veen et al, 2001</a>; <a href="https://doi.org/10.1088/1748-9326/8/3/034021" target="_blank" rel="noopener">Daron and Stainforth, 2013</a>). To run these simulations, we used the Low-EFFourth ensemble generator (<a href="https://doi.org/10.48550/arXiv.2506.03313" target="_blank" rel="noopener">de Melo Vir&iacute;ssimo, 2025a</a>; <a href="https://doi.org/10.5281/zenodo.15566109" target="_blank" rel="noopener">de Melo Vir&iacute;ssimo, 2025b</a>), which is a MATLAB-based framework that allows for large ensembles of low-dimensional dynamical systems to be run and studied in a systematic way (<a href="https://doi.org/10.5194/egusphere-egu23-14755" target="_blank" rel="noopener">de Melo Vir&iacute;ssimo and Stainforth, 2023</a>).</p> <p>These model outputs are presented and discussed in the manuscript "<em>Micro and macro parametric uncertainty in climate change prediction: a large ensemble perspective</em>", published by the Bulletin of the American Meteorological Society (<a href="https://doi.org/10.1175/BAMS-D-24-0064.1" target="_blank" rel="noopener">de Melo Vir&iacute;ssimo and Stainforth, 2025</a>). The manuscript describes the experiments performed, the parameter values used and the modifications done to the original L84-S61 model. For this matter, we also refer you to <a href="https://doi.org/10.1088/1748-9326/8/3/034021" target="_blank" rel="noopener">Daron and Stainforth (2013)</a> and <a href="https://doi.org/10.1063/5.0180870" target="_blank" rel="noopener">de Melo Vir&iacute;ssimo et al. (2024)</a>.</p> <p>All files uploaded were generated from simulations run by the lead author.</p> <p>For specific information about each file uploaded, please refer to the README file. The details of each experiment are also presented in the supplementary materials of the manuscript. If you have any questions, please feel free to contact me.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

A large-ensemble simulation of yields and meteorological drivers to evaluate spatial compounding crop failures in Europe

<p>The dataset consists of a subset from Vogel et al. (2021), comprising large-ensemble simulations of winter wheat yields aggregated at the country level for 20 European countries. The winter wheat yields were simulated by the APSIM-Wheat model (version 7.10) (Zhang et al. 2014) driven by meteorological data from the EC-Earth global climate model (Hazeleger et al., 2010; Van der Wiel et al., 2019). To investigate meteorological drivers of crop failure, the dataset also includes monthly means of daily precipitation, vapour pressure deficit, and maximum temperature fields, of two leading European producers, i.e.&nbsp; France and Germany. For more details, see the description in Vogel&nbsp;et&nbsp;al. (2021).</p> <p>By using this data, you also agree to cite the reference below:</p> <p>Vogel, J., Rivoire, P., Deidda, C., Rahimi, L., Sauter, C. A., Tschumi, E., van der Wiel, K., Zhang, T., Zscheischler, J. (2021). Identifying meteorological drivers of extreme impacts: an application to simulated crop yields. Earth System Dynamics,12(1),151-172.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Synthetic large ensembles from four observation-based products of sea-air CO2 flux from Olivarez et al. (2021)

<p>Synthetic large ensembles from four observation-based products of sea-air CO2 flux from Olivarez et al. (2021). Observation-based products are:</p> <p>Council for Scientific and Industrial Research-Machine Learning (CSIR-ML6)<br> Max Planck Institute Self-Organizing Map-Feed-Forward Neural Network (MPI-SOMFFN)<br> Jena, Germany-Max Planck Institute for Biogeochemistry-Mixed Layer Scheme (JENA-MLS)<br> Copernicus Marine Environment Monitoring Service Feed-Forward Neural Network (CMEMS-FFNN)</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

KNMI-LENTIS large ensemble time slice dataset description

<p><strong>1. Contents&nbsp;</strong></p> <ul> <li><strong>Available variables in KNMI-LENTIS</strong> <ul> <li>request-overview-CMIP-historical-including-EC-EARTH-AOGCM-preferences.txt</li> </ul> </li> <li><strong>Where is the data deposited on the ECWMF&#39;s tape storage (section 4)</strong> <ul> <li>LENTIS_on_ECFS.zip&nbsp;</li> </ul> </li> <li><strong>Data of all variables for 1 year for 1 ensemble member (section 5)</strong> <ul> <li>tree_of_files_one_member_all_data.txt</li> <li>{AERmon,Amon,Emon,LImon,Lmon,Ofx,Omon,SImon,fx,Eday,Oday,day,CFday,3hr,6hrPlev,6hrPlevPt}.zip</li> </ul> </li> </ul> <p><strong>2. Description of this Zenodo dataset</strong></p> <p>This Zenodo dataset pertains to the full KNMI-LENTIS&nbsp;dataset: a large ensemble of&nbsp;simulations with the Global Climate Model&nbsp;EC-Earth3. The periods are for the present-day period (2000-2009) and a future +2K period (2075-2084 following SSP2-4.5). KNMI-LENTIS has 1600 simulated years for both the two climates. This level of sampled climate variability allows for robust and in-depth research into extreme events.&nbsp;The available variables are listed in the file&nbsp;<strong>request-overview-CMIP-historical-including-EC-EARTH-AOGCM-preferences.txt.&nbsp;</strong>All variables are&nbsp;cmorised following CMIP6 data format convention. Further details on the variables and&nbsp;their output dimensions is available via the<a href="https://clipc-services.ceda.ac.uk/dreq/mipVars.html">&nbsp;following search tool</a>.&nbsp;The total size of KNMI-LENTIS&nbsp;is 128 TB. KNMI-LENTIS is stored at the&nbsp;<a href="https://www.ecmwf.int/en/computing/our-facilities/data-handling-system">high performance storage system of the ECMWF (ECFS)</a>.&nbsp;</p> <p>The Global Climate Model that is used for generating this Large Ensemble is EC-Earth3 - VAREX project branch <a href="https://svn.ec-earth.org/ecearth3/branches/projects/varex">https://svn.ec-earth.org/ecearth3/branches/projects/varex</a>&nbsp;(access restricted to ECMWF members).</p> <p>The goal of this Zenodo dataset is :</p> <ol> <li>to&nbsp;provide an accurate description and example of how the KNMI-LENTIS&nbsp;dataset is organised.&nbsp;</li> <li>to&nbsp;describe in which servers the data are deposited and how to gain access to the data for future users</li> <li>to provide links to related git repositories and other content relating to the KNMI-LENTIS production</li> </ol> <p><strong>3. How KNMI-LENTIS is&nbsp;</strong><strong>organised</strong></p> <p>KNMI-LENTIS consists of 2 times 160 runs of 10 years.&nbsp;All simulations have a unique ensemble member label that reflects the forcing, and how the initial conditions are generated. The initial conditions have two aspects: the parent simulation from which the run is branched&nbsp;(macro perturbation, there are 16), and the seed relating to&nbsp;a particular micro-perturbation in the initial three-dimensional atmosphere temperature field (there are 10). The ensemble member label thus is a combination of:&nbsp;</p> <ul> <li>forcing (<em>h</em>&nbsp;for present-day/historical and&nbsp;<em>s</em>&nbsp;for +2K/SSP2-4.5)</li> <li>parent ID (number between 1 and 16)</li> <li>micro perturbation ID (number between 0 and 9)</li> </ul> <p>In this Zenodo dataset we publish 1 year from 1 member to give&nbsp;insight into&nbsp;the type of data and metadata that is representative of the full KNMI-LENTIS&nbsp;dataset. The published data&nbsp;is year 2000 from member&nbsp;<em>h010</em>. See Section 4&nbsp;</p> <p>Further, all KNMI-LENTIS simulations are labeled per the CMIP6 convention of variant labelling. &nbsp;A variant label is made from four components: the realization index&nbsp;<em>r</em>, the initialization index&nbsp;<em>i</em>, the physics index&nbsp;<em>p</em>&nbsp;and the forcing index&nbsp;<em>f</em>. Further details on CMIP6 variant labelling be found in&nbsp;<a href="https://pcmdi.llnl.gov/CMIP6/Guide/modelers.html">The CMIP6 Participation Guidance for Modelers</a>.&nbsp;In the KNMI-LENTIS data set, the forcing is reflected in the first digit of the realization index&nbsp;<em>r&nbsp;</em>of the variant label. For the historical simulations, the one thousands (r1000-r1999) have been reserved. For the SSP2-4.5 the five thousands (r5000-r5999) have been reserved.&nbsp;The parent is reflected in the second and third digit of the realization index&nbsp;<em>r</em>&nbsp;of the variant label (r?01?-r?16?). The seed is reflected in the fourth digit of the realization index&nbsp;<em>r</em>: (r???0-r???9).&nbsp;The seed is also reflected in the initialization index&nbsp;<em>i</em>&nbsp;of the variant label (i0-i9), so this is double information. The physics index&nbsp;<em>p5</em>&nbsp;has been reserved for the ECE3p5 version:&nbsp;all KNMI-LENTIS simulations have the&nbsp;<em>p5</em>&nbsp;label. The forcing index&nbsp;<em>f&nbsp;</em>of the variant label is kept at 1 for all KNMI-LENTIS simulations.&nbsp;As an example, variant label r5119i9p5f1 refers to: the 2K time slice with parent 11 and randomizing seed number 9. The physics index is 5, meaning the run is done with the ECE3p5 version of EC-Earth3.&nbsp;</p> <p><strong>4. Where is the data deposited on the ECWMF&#39;s tape storage</strong></p> <p>In this Zenodo folder, there are several text files and several netcdf files. The text files provide</p> <p>Data from KNMI-LENTIS is deposited in the&nbsp;<a href="https://www.ecmwf.int/en/computing/our-facilities/data-handling-system">ECMWF ECFS tape storage system</a>. Data can be freely downloaded by&nbsp;to those who have access to the ECMWF ECFS. Else, the data can be&nbsp;made available by the authors upon request.&nbsp;</p> <p>The way the dataset is organised is detailed in&nbsp;<strong>LENTIS_on_ECFS.zip.&nbsp;</strong>This contains&nbsp;details on all available KNMI-LENTIS files, in particular&nbsp;details for how these are filed&nbsp;in ECFS.&nbsp;The files on ECFS are tar zipped per ensemble member &amp;&nbsp;variable: these&nbsp;contain 10 years of ensemble member data (10 separate netcdf files). The location on ECFS of the tar-zipped&nbsp;files that are listed in the various text files in this Zenodo dataset is</p> <p>ec:/nklm/LENTIS/ec-earth/cmorised_by_var/&nbsp;</p> <pre><code class="language-bash">#!/bin/bash #------------------- # script to write out LENTIS details on ECFS #------------------- for freq in AERmon Amon Emon LImon Lmon Ofx Omon SImon fx Eday Oday day CFday 3hr 6hrPlev 6hrPlevPt; do for scen in hxxx sxxx; do els -l ec:/nklm/LENTIS/ec-earth/cmorised_by_var/${scen}/${freq}/* &gt;&gt; LENTIS_on_ECFS_${scen}_${freq}.txt done done</code></pre> <p>Further, part of&nbsp;the data will be made publicly available from the&nbsp;<a href="https://esgf-node.llnl.gov/projects/cmip6/">Earth System Grid Federation (ESGF) data portal.</a>&nbsp;We aim to upload&nbsp;most of the monthly variables for the full ensemble. As search terms use&nbsp;<strong>EC-Earth</strong>&nbsp;for model and&nbsp;<strong>p5</strong>&nbsp;for physical index to locate the KNMI-LENTIS data.&nbsp;</p> <p><strong>5. Data of all variables for 1 year for 1 ensemble member</strong></p> <p>The netcdf files of the data of&nbsp;1 year from 1 member&nbsp;<em>h010 </em>are published here&nbsp;to give&nbsp;insight into&nbsp;the type of data and metadata that is representative of the full KNMI-LENTIS&nbsp;dataset. The data are in zipped folders per output frequencies:&nbsp;AERmon, Amon, Emon, LImon, Lmon, Ofx, Omon, SImon, fx, Eday, Oday, day, CFday, 3hr, 6hrPlev, 6hrPlevPt.&nbsp;&nbsp;The&nbsp;text file&nbsp;<strong>request-overview-CMIP-historical-including-EC-EARTH-AOGCM-preferences.txt&nbsp;</strong>&nbsp;gives an overview of variables available per output frequency.&nbsp;&nbsp;the text files&nbsp;<strong>tree_of_files_one_member_all_data.txt&nbsp;</strong>gives an overview of the files in the zipped folders.&nbsp;</p> <p><strong>6. Related links</strong></p> <p>The production of the KNMI-LENTIS ensemble was funded by the KNMI (Royal Dutch Meteorological Institute)&nbsp;multi-year strategic research fund&nbsp;<a href="https://www.knmi.nl/research/weather-climate-models/projects/mso-climate-variability-and-extremes-varex">KNMI MSO Climate Variability And Extremes (VAREX)</a></p> <p>GitHub repository corresponding to this Zenodo dataset:&nbsp;<a href="https://github.com/lmuntjewerf/KNMI-LENTIS_dataset_description.git">https://github.com/lmuntjewerf/KNMI-LENTIS_dataset_description.git&nbsp;</a></p> <p>Github repository for KNMI-LENTIS production code:&nbsp;<a href="https://github.com/lmuntjewerf/KNMI-LENTIS_production_script_train.git">https://github.com/lmuntjewerf/KNMI-LENTIS_production_script_train.git</a></p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Data of the paper: Ensemble daily simulations for elucidating cloud–aerosol interactions under a large spread of realistic environmental conditions

<p>Data of the paper: Ensemble daily simulations for elucidating cloud&ndash;aerosol interactions under a large spread of realistic environmental conditions</p> <p>&nbsp;</p> <p>The name of variable are as in the paper</p>

opencc-by-4.0May 2020View details →
zenodo36/100

Data for manuscript "Modes of Variability in E3SM and CESM Large Ensembles"

An adequate characterization of internal modes of climate variability (MoV) is prerequisite for both accurate seasonal predictions and the attribution and detection of forced climate change in nature. Assessing the fidelity of climate models in simulating MoV is therefore essential; however, doing so is complicated by the large intrinsic variations in MoV and the limited span of the observational record. Large ensembles (LEs) provide a unique opportunity to assess model fidelity in simulating MoV and quantify inter-model contrasts. In this work, these goals are pursued in four recently produced LEs: the Energy Exascale Earth System Model (E3SM) versions 1 and 2 LEs, and the Community Earth System Model (CESM) versions 1 and 2 LEs. In general, the representation of MoV is found to improve across successive E3SM and CESM versions concurrent with improved simulation of the base state climate. The patterns of global coupled modes and many extratropical modes are well simulated by both E3SM2 and CESM2, though various persistent shortcomings are identified. The results both demonstrate the successes of these recent model versions and suggest the potential for continued improvement in the representation of MoV with advances in model physics.

opencc-by-4.0Dec 2022View details →
zenodo36/100

Large ensemble simulations of Holocene temperature and volcanic forcing

<p>Simulations of the global monthly mean volcanic Stratospheric Aerosol Optical Depth (gmSAOD) for 6755 BCE - 1900 CE performed with EVA_H (code available from&nbsp;<span><a href="https://github.com/thomasaubry/EVA_H">https://github.com/thomasaubry/EVA_H</a></span>) and associated Effective Radiative Forcing (ERF).</p> <p>Simulations of the Holocene global annual mean temperature for 6755 BCE - 1900 CE performed with FaIR (code available from <span><a href="https://github.com/OMS-NetZero/FAIR/tree/v2.1.4">https://github.com/OMS-NetZero/FAIR/tree/v2.1.4</a></span>) using volcanic, greenhouse gases (CO2, CH4, N2O), solar, orbital, ice sheets, and anthropogenic land use forcings.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

EnGRaiN : A Supervised Ensemble Learning Method for Recovery of Large-scale Gene Regulatory Networks

<p>EnGRaiN is a supervised machine learning method to construct ensemble networks. To benefit from the typical accuracy advantages of supervised learning methods while taking into account the impossibility of knowing true networks for training, we devised a method that uses small training datasets of true positives and true negatives among gene pairs.</p> <p>The datasets used to evaluate the performance of EnGaiN include (i) simulated datasets generated from Yeast networks and (ii) A. thaliana gene expression datasets.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Float+SOCAT sampling masks for ML reconstruction of surface ocean pCO2 using the Large Ensemble Testbed

<p>Here we provide sampling masks used in the study "The importance of adding unbiased Argo observations to the ocean carbon observing system" (Heimdal &amp; McKinley, 2024, Scientific Reports). In this paper, we reconstruct surface ocean pCO2 using the Large Ensemble Testbed (Gloege et al., 2021, https://doi.org/10.1029/2020GB006788) and the pCO2-Residual method (Bennington et al., 2022, https://doi.org/10.1029/2021MS002960). We provide 2 different sampling masks used in the experiments presented in Heimdal &amp; McKinley (2024). These masks represent two different float sampling schemes (+SOCAT) including 500 floats, corresponding to historical Argo float observations (https://fleetmonitoring.euro-argo.eu/dashboardpatterns) and potential optimized float sampling (following Chamberlain et al., 2023, <a href="https://doi.org/10.1175/JTECH-D-22-0093.1" target="_blank" rel="noopener">https://doi.org/10.1175/JTECH-D-22-0093.1</a>).&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Large ensemble climate modelling time series for the Rhine catchment, including drought2018 storylines

<p>Dataset associated with&nbsp;<strong>Van der Wiel, Lenderink, De Vries (2021):&nbsp;Physical storylines of future European drought events like 2018 based on ensemble climate modelling,&nbsp;<em>Weather and Climate Extremes, </em>DOI <a href="http://doi.org/10.1016/j.wace.2021.100350">10.1016/j.wace.2021.100350</a>.</strong></p> <p>Large ensemble climate modelling time series for the Rhine catchment. The dataset contains three ensembles (present-day, pre-industrial + 2C-warming, pre-industrial +&nbsp;3C-warming) of 2000 years each, various variables related to drought are included.&nbsp;All data is derived from the EC-Earth global climate model (v2, Hazeleger et al. 2012, DOI <a href="https://doi.org/10.1007/s00382-011-1228-5">10.1007/s00382-011-1228-5</a>). Descriptions of large ensemble experimental setup can be found in Van der Wiel et al. (2019, DOI <a href="http://doi.org/10.1029/2019GL081967">10.1029/2019GL081967</a>). Files: *_d_ECEarth_??_Rhine.tar.gz</p> <p>Additionally, three sets of storylines of droughts similar to the western European drought of 2018 are included.&nbsp;These are the simulated events selected from the large ensembles, using metrics 1, 2 and 3 of Van der Wiel et al. (2021, DOI <a href="http://doi.org/10.1016/j.wace.2021.100350">10.1016/j.wace.2021.100350</a>). Files:&nbsp;drought18_m[123]_Rhine.tar.gz</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

GFDL-FLOR Large Ensemble Arctic Sea Ice Data

<p>This upload contains Arctic sea ice data from the GFDL-FLOR Large Ensemble and related data analysis code, as published in Bushuk et al. (2020). See readme.txt for a description of the datasets and code.</p> <p>Reference: Bushuk, M., M. Winton, D. Bonan, E. Blanchard-Wrigglesworth, T. Delworth, 2020: A mechanism for the Arctic sea ice spring predictability barrier, Geophysical Research Letters, in press.</p>

opencc-by-4.0May 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record