Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
138
datasets available to search
ShareScore release 0.9.0
Dataset results
138 results for “emulator”
Multi-fidelity Gaussian Process Emulation for Atmospheric Radiative Transfer Models
<p>This repository contains several datasets of spectral atmospheric transfer functions (i.e. path radiance, transmittances, spherical albedo) simulated with MODTRAN6 atmospheric radiative transfer model. The simulations are stored in hdf5 files using the Atmospheric Look-up table Generator (ALG) toolbox (<a href="https://doi.org/10.5194/gmd-13-1945-2020">https://doi.org/10.5194/gmd-13-1945-2020</a>). Each dataset has an associated .xml file that includes the configuration of ALG/MODTRAN6 executions. All datasets include the input atmospheric/geometric variables that are summarized in the following table. Each dataset file has a random distribution (based on latin hypercube sampling) these input variables with varying number of points (e.g. train500.h5 contains 500 samples). The <em>reference </em>dataset contains 10000 samples and was used as reference for evaluating Gaussian Processes emulators.</p> <table> <tbody><tr> <th>Input Variables</th> <th>Units</th> <th>Min</th> <th>Max</th> </tr> </tbody><tbody> <tr> <td>O3 column concentration</td> <td>atm-cm</td> <td>0.25</td> <td>0.45</td> </tr> <tr> <td>Columnar Water Vapor</td> <td>g/cm2</td> <td>0.2</td> <td>4</td> </tr> <tr> <td>Aerosol Optical Thickness</td> <td>-</td> <td>0.04</td> <td>0.6</td> </tr> <tr> <td>Asymmetry parameter</td> <td>-</td> <td>0.5</td> <td>0.85</td> </tr> <tr> <td>Angstrom exponent</td> <td>-</td> <td>0.1</td> <td>2</td> </tr> <tr> <td>Single Scattering Albedo</td> <td>-</td> <td>0.8</td> <td>1</td> </tr> <tr> <td>Surface elevation</td> <td>km</td> <td>0</td> <td>2.5</td> </tr> <tr> <td>Solar Zenith Angle</td> <td>deg</td> <td>0</td> <td>70</td> </tr> <tr> <td>Relative Zenith Angle</td> <td>deg</td> <td>0</td> <td>180</td> </tr> </tbody> </table> <p> </p>
UKESM1.0-ice simulation output used as test data in Burgard et al., Emulating present and future simulations of melt rates at the base of Antarctic ice shelves with neural networks
<p>These files contain NEMO ocean model output and domain definitions for the Southern Ocean from UKESM1.0-ice simulations described in section 6.3.2 of Smith et al. "Coupling the U.K. Earth System Model to Dynamic Models of the Greenland and Antarctic Ice Sheets" , Journal of Advances in Modeling Earth Systems, 2021</p> <p>Files labelled "bf663" are the UKESM simulation referred to in that section as "constant 1970 greenhouse gas and other forcings". Files labelled bi646 are the UKESM simulation referred to in that section as "instantaneously quadrupled 1970 CO<sub>2</sub> concentrations".</p> <p>They were used as test data for the performance of neural networks in Burgard et al., "Emulating present and future simulations of melt rates at the base of Antarctic ice shelves with neural networks", Journal of Advances in Modeling Earth Systems 2023.</p>
GF4ACE -- Data from: Reanalysis-based global radiative response to sea surface temperature patterns: Evaluating the Ai2 climate emulator
Open the record for dataset details and reuse information.
Emulator-based decomposition for structural sensitivity of core-level spectra
Open the record for dataset details and reuse information.
AgMIP's GGCMI Phase II: Crop model Emulators at 0.5 degree global resolution
<p>GGCMI Phase II crop model polynomial emulator parameters. See https://www.geosci-model-dev.net/special_issue1062.html for more details. A0: no growing season adaptation. A1: with growing season adaptation.</p>
Data for Gaussian-Process-Based Emulators for Building Performance Simulation
<p>The ZIP folder contains the MAT files you need to rerun the experiment described in</p> <p>Rastogi, Parag, Mohammad Emtiyaz Khan, and Marilyne Andersen. 2017. “<strong>Gaussian-Process-Based Emulators for Building Performance Simulation</strong>.” In <em>Proceedings of BS 2017</em>. San Francisco, CA, USA: IBPSA.</p> <p>-------------------------------------------------------------</p> <p>The two m-scripts (MATLAB) help you to load the results reported in the paper. Make sure to CHECK the file paths inside the scripts, especially to the MAT files. Usually, the paths should be fine if you update the variable <em>pathMATfolder</em> inside the script <em>RunThis.m</em> .</p> <p>There are two types of MAT files inside the folder called "Data" :</p> <p>1. Original data (building simulations) --> gpdata_BaseSimulation.mat</p> <p>2. Errors and predictions - errs_BaseSimulation_N_M.mat and ystore_BaseSimulation_N_M.mat --> The first contains all the error quantities and the second the 'y' predictions. The number N represents the run number (subset of master training data set sampled for the given run). The number M can take only two values - 1 or 2. The models for heating load are represented by 1 and for cooling by 2.</p> <p>3. Metadata - trainN* --> These files contain metadata for setting up the plots.</p> <p>See github repository <strong>https://github.com/paragrastogi/GPregressionInBS.git</strong> for more scripts. See <strong>www.paragrastogi.com</strong> or <strong>www.ibpsa.org</strong> for the conference paper.</p>
Model checkpoints for "SEEDS: Emulation of Weather Forecast Ensembles with Diffusion Models"
<p>Checkpoints for all SEEDS models in the paper <a href="https://arxiv.org/abs/2306.14066" rel="nofollow">https://arxiv.org/abs/2306.14066</a>, including all SEEDS-GEE and SEEDS-GPP models in the main manuscript, and the additional models trained in the Supplemental Material.</p> <p>Checkpoint naming convention:</p> <ul> <li><code>gee_c2_s7</code>: SEEDS-GEE trained conditioning on 2 seeds for 7-day lead time.</li> <li><code>gpp_c2_s7_g3_r4</code>: SEEDS-GPP trained conditioning on 2 seeds for 7-day leadtime, where the label mixture is 3 GEFS members and 4 ERA reanalyses.</li> </ul>
Data and code for firn air content emulator
<p>This data corresponds to the paper: "Future (2015-2100) Antarctic-wide ice-shelf firn air depletion from a statistical firn emulator".</p> <p>In this new version we share the data and code used in this manuscript.</p> <p>Emulators.zip contains the models used to calculate FAC. The ‘weights’ files are used to standardize data based on the training dataset.</p> <p>fac_tables.zip contains csv files of CMIP6 model min, median, and max FAC for each scenario for each ice shelf.</p> <p>masks.zip contains ice shelf masks and other useful files used in the analysis</p> <p>figures.zip contains code and data used to make figures</p> <p>code.zip contains code used to create and run emulator</p> <p>SNOWPACK_FAC.zip contains FAC results from the SNOWPACK model as netcdf files for each emission scenario</p> <p>ERA5.zip contains the ERA5 climate output used for this analysis</p> <p>ssp***<em>_fac_*</em>.nc.zip are the FAC emulator results using the TP1 ice lens thickness-permeability relationship</p> <p>ssp***_allCMIP6.nc.zip is the CMIP6 climate output from each of the 34 models used in this study. </p>
Processed dataset used for developing tsunami inundation emulators(Part-III Simulation and Emulation Dataset)
<div> <div>This dataset is related to the main Zenodo repository: https://doi.org/10.5281/zenodo.13738078</div> <br> <div>This dataset contains some of the processed datasets covering tsunami inundation depths using simulation results and emulation results discussed in the article - "Towards Using Machine Learning Emulation for Probabilistic Inundation Mapping: Multiple Earthquake Sources and Near-field Effectwith project repo - https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators.</div> <br> <div>The post-processed (numpy) files for the two test locations of Catania(CT) and Siracusa(SR) are provided in compressed gzip files(.gz):</div> <br> <div>Processed inundation depth files from simulation and emulation typically stored in model/CT/ptha or model/SR/ptha</div> <div> -<strong>pthaSR.tar.gz </strong>- max inundation depth1d(eve x m) for true(simulation and pred(emulation of different models))</div> <div> -<strong>pthaCT.tar.gz</strong> - max inundation depth1d(eve x m) for true(simulation and pred(emulation of different models))</div> <br> <div>The filename follows the nomenclature with:</div> -true - simulation results</div> <div>-pred - emulator with pretraining</div> <div>-pred_direct - emulator without pretraining</div> <div>-pred_nodeform - emulator without pertaining and deformation input</div> <div><br> <div> <div>More information on the attached readme, see project structure and code is available at:</div> <div><strong>https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators</strong></div> </div> </div> <p> </p>
Processed files used for OQ based tsunami loss modelling using HPC based inundation and emulators(Part-IV OQ Simulation and Emulation Hazard Dataset)
<div> <div>This dataset is related to the main Zenodo repository: https://doi.org/10.5281/zenodo.13738078</div> <br> <div>This dataset contains some of the processed OpenQuake files covering tsunami inundation depth hazard using simulation results and emulation results discussed in the article - "Towards Using Machine Learning Emulation for Probabilistic Inundation Mapping: Multiple Earthquake Sources and Near-field Effects with project repo - https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators.</div> <div> </div> <div>The risk calculation and procedure is available in the main repo: <a href="https://github.com/naveenragur/ML4SicilyTsunami/tree/main-dev/risk">https://github.com/naveenragur/ML4SicilyTsunami/tree/main-dev/risk </a></div> <div> </div> <br> <div>The processed files for the test locations of Catania(CT) are provided in compressed gzip files(.gz):</div> <br> <div>Processed hazard information used to prepare and run OQ event based risk analysis are as below,</div> <div> -<strong>hazard.tar.gz </strong>- preliminary numpy files with sitcol, eventid and hazard magnitude info</div> <div> -<strong>loss.tar.gz</strong> -final hdf5 files with both event hazard and site info</div> <br> <div>The filename follows the nomenclature with:</div> </div> <div>ML4SicilyTsunami/risk/loss/tsunami_prob_892.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_1658.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_3454.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_7071.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_true.hdf5</div> <div><br> <div> <div>More information on the attached readme, see project structure and code is available at:</div> <div><strong>https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators</strong></div> <div><strong>https://github.com/naveenragur/ML4SicilyTsunami/tree/main-dev/risk</strong></div> <div><strong>https://github.com/naveenragur/OQ-Tsunami</strong></div> </div> </div>
Met Office UKCP Local CPM precipitation ML emulator dataset
<div> <div> <div> <h1>Met Office UKCP Local CPM precipitation ML emulator dataset</h1> <p>This is a collection of two datasets: one sourced from CPM data (bham64_ccpm-4x_12em_psl-sphum4th-temp4th-vort4th_pr.tar.gz) and one sourced from GCM data (bham64_gcm-4x_12em_psl-sphum4th-temp4th-vort4th_pr.tar.gz). Each dataset is made up of climate model variables extracted from the Met Office's storage system, combining many variables over many years. It consists of 3 NetCDF files (train.nc, test.nc and val.nc), a YML ds-config.yml file and a README (similar to this one but tailored to the source of the data). Code used to create the dataset can be found here: <a href="https://github.com/henryaddison/mlde-data">https://github.com/henryaddison/mlde-data</a> (specifically the james-submission tag).</p> <p>The YML file contains the configuration for the creation of the dataset, including the variables, scenario, ensemble members, spatial domain and resolution, and the scheme for splitting the data across the three subsets.</p> <p>Each NetCDF contains the same variables but split into different subsets (train, val and test) of the based on time dimension.</p> <p>Otherwise the NetCDF files have the sames dimensions and coordinates for ensemble_member, grid_longitude and grid_latitude.</p> <ul> <li>Spatial resolution: This has two parts - the resolution of the data and the grid resolution stored at in the file. For predictand variables this is 2.2km variables coarsened 4 times to 8.8km (this is the target grid). For predictor variables this is 2.2km variables conservatively regriddded to GCM 60km grid or variables from GCM (so already on 60km grid) then regrid (nearest neighbour) to the target grid of predictands. In the naming convention of resolution used in config files, 60km resolution is synonamous with the GCM grid and 2.2km resolution is synonamous with the CPM grid.</li> <li>Spatial domain: A 64x64 section of the 8.8km target grid covering England and Wales</li> <li>Time resolution: daily</li> <li>Time domain: 1st Dec 1980 to 30th Nov 2000; 1st Dec 2020 to 30th Nov 2040; 1st Dec 2060 to 30th Nov 2080. Uses a 360-day calendar.</li> <li>Scenario: RCP8.5</li> <li>Ensemble Members: 01, 04-13 & 15 (these correspond to the 12 ensemble member runs from the CPM but don't carry intrinsic meaning).</li> <li>Split scheme: 70% training, 15% validation, 15% testing, split by choosing complete seasons at random, with an equal number of each season from each of the 3 time periods.</li> </ul> <p> </p> <h2>Predictor variables</h2> <ul> <li>psl (hPa) - mean sea level pressure</li> <li>temp850, temp700, temp500, temp250 - air temperature (K) at 850, 700, 500 and 250 hPa</li> <li>vorticity850, vorticity700, vorticity500, vorticity250 - relative vorticity (s^-1) at 850, 700, 500 and 250 hPa</li> <li>spechum850, spechum700, spechum500, spechum250 - specific humidity at 850, 700, 500 and 250 hPa</li> </ul> <h2>Predictand variable</h2> <ul> <li>target_pr - precipitation rate (mm/day)</li> </ul> <p> </p> <p>UPDATE 2025-03-27: Dataset tars are renamed to make it clearer their source (ccpm for coarsened CPM and gcm for GCM).</p> </div> </div> </div>
Data archive for paper "Machine Learning Emulation of Urban Land Surface Processes"
<p>This archive contains models, data* (Overview), as well as the Singularity image to optionally rerun experiments described in "<a href="https://doi.org/10.1029/2021MS002744">Machine Learning Emulation of Urban Land Surface Processes</a>".</p> <p><strong>Prerequisites</strong></p> <ul> <li>Linux or macOS with Bash shell.</li> <li><a href="https://sylabs.io/">Singularity</a> (tested with version 3.6.3-1.el8)</li> </ul> <p>Please note that all steps require <a href="https://sylabs.io/">Singularity</a> to be installed on your system. If you are looking for information on how to install or use Singularity, please refer to the <a href="https://sylabs.io/docs">Singularity documentation</a>.</p> <p><strong>Overview</strong></p> <p>A general overview of the repository structure is given below. Due to licensing restrictions analysis and forcing data (*) cannot be included and need to be requested separately (see Initialization). Data derivatives (**) from either analysis or forcing, as well as intermediary data (***), are not included as they can be generated by rerunning experiments (see Usage).</p> <pre><code>. ├── data │ ├── analysis* │ ├── forcing* │ ├── teb │ ├── utils │ ├── wps │ └── wrf ├── hpc ├── models │ ├── teb │ ├── unn │ ├── wps │ └── wrf-unn ├── notebooks ├── outputs │ ├── analysis** │ ├── benchmark*** │ ├── forcing** │ ├── kerastuner*** │ ├── notebooks │ ├── tabular │ ├── teb** │ ├── unn** │ ├── wps*** │ └── wrf ├── paper │ └── figures ├── singularity └── tools </code></pre> <p><strong>Initialization</strong></p> <p>Forcing and analysis data need to be requested separately. The following directories should map to their respective data archives:</p> <ul> <li><code>./data/analysis</code> -> <a href="http://doi.org/10.5281/zenodo.4678387">Grimmond et al. (2013)</a></li> <li><code>./data/forcing</code> -> <a href="http://doi.org/10.5281/zenodo.4679279">Grimmond et al. (2021)</a></li> </ul> <p><strong>Usage</strong></p> <p>To rerun all experiments and reproduce results, run <code>tools/run_all.sh</code> from your command prompt. After completion, all results are saved in the <code>outputs</code> directory. Note that WRF simulations require high CPU time and may take hours or days to complete.</p> <p>Alternatively, if <a href="https://en.wikipedia.org/wiki/Portable_Batch_System">Portable Batch System (PBS)</a> is available on your system, the following helpers may be used instead:</p> <pre><code>qsub hpc/submit_init.pbs qsub hpc/submit_tuner.pbs qsub hpc/submit_unn.pbs qsub hpc/submit_find_median_unn.pbs qsub hpc/submit_wrf.pbs qsub hpc/submit_postprocess.pbs qsub hpc/submit_benchmark.pbs </code></pre> <p>Note that you may need to modify PBS helper scripts to suit your specific environment.</p> <p><strong>Development notes</strong></p> <p>See DEVELOP.md.</p> <p><strong>License</strong></p> <p>The source code developed for this work is licensed under MIT (<code>LICENSE_CODE.txt</code>). For licensing information of third-party software see licenses under the <code>models</code> directory. Data files in this archive, including the initial and boundary condition data from the European Centre for Medium-Range Weather Forecasts (<code>data/wps/ungrib</code>), are licensed under CC BY-NC 4.0 (<code>LICENSE_DATA.txt</code>).</p>
Training data and models for microphysics emulation
<p>Training data and models for microphysics emulation</p> <p>The training data is a subset of the full dataset described in the NeurIPS submission. Roughly speaking, 30 day runs with FV3GFS, with zhao carr microphysics. To keep the data reasonable in size, 1000 random netCDFs are sampled from the over 7000 files in the full training dataset. 200 test files are sampled.</p> <p>Also contains the trained ML models at models/</p> <p>Data behind the plots and tables is at plot-data/.</p> <p> </p>
Graph neural network emulator for modeling of ice dynamics and calving in the Helheim Glacier, Greenland
<p>These files include the following codes and datasets for developing graph neural network (GNN) emulators for the Ice-sheet and Sea-level System Model (ISSM) for modeling ice sheet dynamics and calving in the Helheim Glacier, Greenland.</p> <ul> <li>ISSM_DGL_Helheim.py: Python file for training GNN models</li> <li>ISSM_CNN_Helheim.py: Python file for training convolutional neural network (CNN) models</li> <li>*.mat: Datasets of the ISSM transient simulation results</li> </ul>
A crop yield change emulator for use in GCAM and similar models: Persephone v1.0
<p>This is an archive of the raw data and analysis source code for the paper "A crop yield change emulator for use in GCAM and similar models: Persephone v1.0". The archive contains:</p> <ul> <li><strong>data.zip:</strong> All source code for analysis, input data for analysis, and results of analysis</li> <li><strong>persephone.proj : </strong>R project for ease of reproducing analysis</li> </ul>
Training and testing data, associated code and estimators for emulating a convection scheme
<p>Data and code for a random-forest convection scheme associated with the paper:</p> <p>"Using machine learning to parameterize moist convection: potential for modeling of climate, climate change and extreme events"</p> <p>by Paul A. O'Gorman and John G. Dwyer (to appear in JAMES)</p>
Data perennial firn aquifer emulator
<div> <p>These NetCDF files contains perennial firn aquifer output from an XGBoost firn emulator for Antarctica. The datasets are time series of annual perennial liquid water content (if present), and otherwise the days per year without the presence of liquid water within the firn (times -1). The dataset covers the period 2015-2100 for several RCM and GCM combinations. Please feel free to send me a message if you have any questions (s.b.m.veldhuijsen@uu.nl). These data are used in a manuscript to be submitted to <strong><em>The Cryosphere</em></strong>.</p> <p> </p> </div>
FAC emulator models and results
<p>This is the emulator results for the paper: "Future (2015-2100) Antarctic-wide ice-shelf firn air depletion from a statistical firn emulator" (in prep). Here you can find the emulator models for each ice lens thickness-permeability relationship used (TP1-3) and netcdf results with ice-shelf FAC for each CMIP6 model.</p>
TARDIS configuration and emulator weights and training data for "1991T-Like Type Ia Supernovae as an Extension of the Normal Population"
<p>This dataset contains two archives of data related to the paper "1991T-Like Type Ia Supernovae as an Extension of the Normal Population"<br> <br> The first dataset, <a href="https://zenodo.org/api/files/de696fe0-3280-44f2-8ef4-975b92fad260/TARDIS_Emulator_Config.tar.gz">TARDIS_Emulator_Config.tar.gz </a>, contains the atomic data used to run TARDIS and a template configuration file from which samples are generated including the flags for the physics implementation used.</p> <p>The second dataset, InferenceScripts.tar.gz, contains the trained probabilistic neural network, the training/validation data (Under NNData), and scripts used to train the model and load and evaluate the model. Scripts that perform inference on spectra, as well as a folder of observed spectra (Under CorrectedSpectra), are included as well. A conda environment yaml file is included to rebuild the Python environment required to run all of the scripts. For questions please email John O'Brien.</p>
Deep Learning Regional Climate Model Emulators: a comparison of two downscaling training frameworks [datasets]
<p>Outputs used in:</p> <p><em>van der Meer, M., de Roda Husman, S., Lhermitte, S.: </em>Deep Learning Regional Climate Model Emulators: a comparison of two downscaling training frameworks</p> <ul> <li>MAR(ACCESS1-3)_monthly_SMB.nc: MAR outputs with monthly values of SMB and components over the Antarctic ice sheet (1980--2100)</li> <li>MAR(ACCESS1-3)-stereographic_monthly_GCM_like.nc: MAR outputs upscaled to GCM resolution (1980--2100)</li> <li>ACCESS1-3-stereographic_monthly_cleaned.nc: GCM monthly outputs over the Antarctic ice sheet (1980--2100)</li> </ul> <p>The up-to-date working versions of our experiments and source code can be found and are available on our GitHub: <a href="https://github.com/marvande/RCM-Emulator">https://github.com/marvande/RCM-Emulator</a> and at this link: <a href="https://doi.org/10.5281/zenodo.7875967">https://doi.org/10.5281/zenodo.7875967</a></p> <p>Data usage notice:</p> <p>If you use any of these results, please acknowledge the work of the people involved in producing them. You should also refer to and cite the following paper:</p> <p><strong>Cite as: </strong>Marijn van der Meer, Sophie de Roda Husman, S Lhermitte. Deep Learning Regional Climate Model Emulators: a comparison of two downscaling training frameworks. <em>Authorea.</em> December 27, 2022 <br> DOI: <a href="https://doi.org/10.22541/essoar.167214210.02213149/v1">10.22541/essoar.167214210.02213149/v1</a> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.