Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
138
datasets available to search
ShareScore release 0.9.0
Dataset results
138 results for “emulation”
Emulator of PR-DNS: Accelerating Dynamical Fields with Neural Operators in Particle-Resolved Direct Numerical Simulation
<p>The codes directory includes the various machine learning models, such as FNO, UNet and ResNet. R128_init1 and R128_init2 are the PR-DNS time step simulations at different initial conditions. R64_init2, R128_init2 and R256_init2 are the PR-DNS time step simulations at different resolutions. </p> <p> </p>
Dataset for "An Accuracy Study of Emulation Daemons for IEEE 802.11 Networks"
<p>This is the dataset used in "An Accuracy Study of Emulation Daemons for IEEE 802.11 Networks" presented at the 48th IEEE Conference on Local Computer Networks (LCN), October 1-5, 2023, Daytona Beach, Florida, USA.</p> <p>`raw_results` contains all the raw results redacted from the experiment hosts, including log files.<br> The `tikz` files used to generate the figures presented in the paper can be found in `tex/figures`. Executing `make` in the `tex` directory generates the `pdf` plots in the `figures` folder.<br> The `figures` folder also contains pre-processed results from the raw data set.</p>
Glacier catalogue for IGM physics-informed deep-learning emulator pretraining
<p>This dataset was created with the iceflow glacier model CfsFlow to generate glacier extent and retreat in the Alps and New zealand with the goal to generate realistic and diverse glacier states for pretraining the physics-informed deep-learning emulator of IGM (https://github.com/jouvetg/igm).</p> <p>The data consists of distributed surface topography (usurf) and ice thickness (thk) of 8 snapshots of 37 glaciers in different stages (advance and retreat). The data is organized glacier-wise: each folder corresponds to one glacier, which contains a unique NetCDF file with 2D distributed raster data of surface elevation and ice thickness.</p>
Emulation of the KEYNOTE-189 Trial Using Electronic Health Records
ClinicalTrials.gov study NCT05908799. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Impact of PSA-based Screening on Mortality Among Men Aged 75-79: Target Trial Emulation
ClinicalTrials.gov study NCT07206693. IPD Sharing: NO. Countries: 1. Publications: 1.
Do harvest retention patches in the boreal forest emulate those resulting from wildfire? A comparison of understory vegetation a decade after disturbance
Open the record for dataset details and reuse information.
GOES-16 ABI and collocated SNPP-VIIRS imagery for evaluating the emulation of daytime cloud products at night
Open the record for dataset details and reuse information.
FD-detector: Dataset gathered running measurements in the wild and Emulations for validation
<p>This file contains two folders: the dataset and emulation setup.</p> <p>The dataset is composed of warts files, resulting from traceroutes run with Scamper on NLNOG RING nodes. On the other hand, the emulations were performed on GNS3.</p>
Dataset associated with "Emulating subglacial hydrology in ice sheet models with deep learning methods" by Verjans and Robel.
<p>See Readme file for descriptions.</p>
Emulating present and future simulations of melt rates at the base of Antarctic ice shelves with neural networks
<p>This dataset contains the data and scripts for the publication "<a href="https://doi.org/10.1029/2023MS003829">Emulating present and future simulations of melt rates at the base of Antarctic ice shelves with neural networks</a>" in <em>Journal of Advances in Modeling Earth Systems</em>.</p> <p>Before going into details, here is a reminder that the NEMO runs for the training dataset are called 'OPM+number'. These are the corresponding names given in the manuscript: OPM006=HIGHGETZ, OPM016=WARMROSS, OPM018=COLDAMU and OPM021=REALISTIC. For the testing dataset: 'bf663' is the REPEAT1970 run and 'bi646' is the 4xCO2 run.</p> <p>Most of the formatting and preprocessing of the training data has been made for <a href="https://tc.copernicus.org/articles/16/4931/2022/">Burgard et al. 2022</a>. The raw and to-some-degree processed data can therefore be found here: <a href="https://doi.org/10.5281/zenodo.7308352">https://doi.org/10.5281/zenodo.7308352</a>. <br>The raw data for the testing dataset is from <a href="https://doi.org/10.1029/2021MS002520">Smith et al. 2021</a>, you can find it here: <a href="doi.org/10.5281/zenodo.7886986">https://doi.org/10.5281/zenodo.7886986</a></p> <p>The following folders and files can be found here:</p> <p>===============<br><strong>raw/</strong></p> <p>Some geometrical files needed for initial data formatting and masking.</p> <p>===============<br><strong>interim/</strong></p> <ul> <li><strong>ANTARCTICA_IS_MASKS</strong>/ (<em>from INTERIM_ANTARCTICA_IS_MASKS.zip</em>): contains <ul> <li>masks and geometric information for the testing dataset to be included in the input file of the neural network and for the classic parameterisations.</li> <li>local bedrock and ice meridional and zonal slopes</li> </ul> </li> <li><strong>BOXES/</strong> (<em>from INTERIM_BOXES.zip</em>): contains variables needed to apply the box parameterisation for the testing dataset</li> <li><strong>PLUMES/ </strong>(<em>from INTERIM_PLUMES.zip</em>): contains the variables needed to apply the plume parameterisation for the testing dataset</li> <li><strong>SMITH_bf663/</strong><em><strong> and </strong></em><strong>SMITH_bi646/ </strong>(<em>from INTERIM_SMITH*.zip</em>): for testing dataset, <ul> <li>corrected_draft_bathy_isf.nc: file containing ice draft and bathymetry corrected by ice draft concentration to account for the biased draft and bathymetry at the grounding line resulting from the interpolation from the native NEMO grid to the stereographic grid (values under ice shelf and NaNs over land</li> <li>custom_lsmask_Ant_stereo_clean.nc: land-sea mask giving 0 = ocean, 1 = shelf, 2 = land</li> <li>isfdraft_conc_Ant_stereo.nc: ice-shelf concentration resulting from the interpolation from the native NEMO grid to the stereographic grid</li> <li>other_mask_vars_Ant_stereo.nc: contains other variables used for the masks</li> <li>the reference melt: 1D containing integrated melt, 2D containing melt fields, box1 containing melt near the grounding line</li> </ul> </li> <li><strong>T_S_PROF/ </strong>(<em>from INTERIM_T_S_PROF.zip</em>) <ul> <li>Mean profiles used as input for traditional parameterisations</li> <li>T and S 2D fields, extrapolated from the mean profiles to the local ice draft depth (needed as input for the neural network)</li> <li>Fields of mean and standard deviation T and S for all points (needed as input for the neural network)</li> </ul> </li> <li><strong>NN_MODELS/</strong><em><strong> </strong>(from </em>INTERIM_NN_MODELS<em>.zip</em>) contains all neural networks trained for this paper (for the cross validation and over the whole dataset for testing)</li> <li><strong>INPUT_DATA/ </strong>(from INTERIM<em>_</em>INPUT_DATA.zip) contains all input csv files containing the input datasets for the different training and testing iterations. Also contains the metrics to normalise the input. For the cross-validation, the input csv files are not included because they are too large. However, they can be reconstructed from the individual files for ice shelves and time blocks. The metrics to normalise the data during the cross-validation are included in EXTRAPOLATED_ISFDRAFT_CHUNKS_CV!</li> </ul> <p>===============<br><strong>processed/MELT_RATE/</strong></p> <p>Contains resulting melt rates</p> <ul> <li><strong>CV_ISF :</strong> Cross-validation results over ice shelves</li> <li><strong>CV_TBLOCKS : </strong>Cross-validation results over time</li> <li><strong>SMITH_bf663 : </strong>Neural network results for REPEAT1970</li> <li><strong>SMITH_bf663_CLASSIC : </strong>"Traditional" parameterisation results for REPEAT1970</li> <li><strong>SMITH_bi646 :</strong> Neural network results for 4xCO2</li> <li><strong>SMITH_bi646_CLASSIC:</strong> "Traditional" parameterisation results for 4xCO2</li> </ul> <p>=====================</p> <p>The explanation around the scripts can be found in README.rst with the scripts in<strong> scripts_paper_simpleNN_basal_melt.zip</strong>.<br><em>Note that these are the scripts needed to produce the results in the paper. You can also find them on Github: </em><a href="https://github.com/ClimateClara/https://github.com/ClimateClara/scripts_paper_simpleNN_basal_melt"><em>https://github.com/ClimateClara/scripts_paper_simpleNN_basal_melt</em></a>, <em>find the most up-to-date version of the package 'multimelt' here: </em><a href="https://github.com/ClimateClara/multimelt"><em>https://github.com/ClimateClara/multimelt</em></a><em> and a version you can install via pip here: </em><a href="https://github.com/ClimateClara/multimelt"><em>https://pypi.org/project/multimelt/</em></a></p> <p>Finally, if anything is unclear, check out the "Methods" section of the paper: <a href="https://doi.org/10.1029/2023MS003829">https://doi.org/10.1029/2023MS003829</a></p>
Processed dataset used for developing tsunami inundation emulators(Part-II Test Dataset)
<div> <div>This dataset is related to the main Zenodo repository: https://doi.org/10.5281/zenodo.13738078</div> <br> <div>This dataset contains some of the processed datasets covering testing datasets covering tsunami inputs and outputs (for parameters of offshore waveforms, local deformation fields and inundation depths) used in the evaluation of machine learning emulator discussed in the preprint article - "Towards Using Machine Learning Emulation for Probabilistic Inundation Mapping: Multiple Earthquake Sources and Near-field Effect with project repo - https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators.</div> <br> <div>The post-processed (numpy) files for the two test locations of Catania(CT) and Siracusa(SR) are provided in compressed gzip files(.gz) typically stored in data/processed:</div> <strong>-test_dZ.tar.gz</strong> (local deformation files)</div> <div><strong>-test_d.tar.gz </strong>(inundation depth files)</div> <div><strong>-test_t.tar.gz</strong> (offshore waveform files)</div> <p>The filename follow nomenclature as below</p> <p><strong>d_CT_0.dat, dZ_CT_0.dat, dflat_CT_0.dat, dZflat_CT_0.dat, t_CT_0.dat,lat_lon_idx_CT_892.npy</strong> where the file names represent <strong>{parameter}_{site}_{size}</strong></p> <ol> <li>The first var represent parameter of the file. <ul> <li>d - 2 dimensional file for inundation depth (<strong>events </strong>x <strong>m </strong>x <strong>n</strong>)</li> <li>dZ - 2 dimensional file for local deformation( <strong>events </strong>x <strong>m</strong> x <strong>n</strong>)</li> <li>dflat - 1 dimensional flat file for inundation ( <strong>events </strong>x <strong>locations</strong>)</li> <li>dZflat - 1 dimensional flat file for local deformation( <strong>events </strong>x <strong>locations</strong>)</li> <li>t - 3 dimension offshore waveform for (<strong>events </strong>x <strong>gauges </strong>x <strong>timesteps</strong>)</li> <li>lat_lon_idx - index file for matching lat long coordinate with location indices(<strong>locations x lat x lon</strong>)</li> </ul> </li> <li>The second var represents site of the file. <ul> <li>CT - Catania</li> <li>SR - Siracusa</li> </ul> </li> <li>The third var represents number of events in the file or the selection size. <ul> <li>0,1,2,3 - represent the approx 50000 test events divided into 4 splits</li> </ul> </li> </ol> <div>More information on the attached readme, see project structure and code is available at:</div> <div><strong>https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators</strong></div>
Processed dataset used for developing tsunami inundation emulators(Part-I Training Dataset and Model Weights)
<p>This dataset contains some of the processed training datasets covering tsunami inputs and outputs parameters (for offshore waveforms, local deformation fields and inundation depths used in developing a machine learning emulator discussed in the preprint article -<strong> Towards Using Machine Learning Emulation for Probabilistic Inundation Mapping: Multiple Earthquake Sources and Near-field Effects</strong> with project repo - <a href="https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators.">https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators.</a></p> <p>####Code####</p> <p>Source code from the above github project : <strong>ML4SicilyTsunami-NPJ2024.tar.gz</strong></p> <p>The post-processed numpy dat files for the two test location of <strong>Catania(CT) and Siracusa(SR)</strong> are provided for each variable retrieved from data/processed folder of the project as compressed gzip files(.gz):</p> <p>####Model Checkpoints#### <br>ML Model checkpoints from all training experiments.<br><strong>- outSR.tar.gz (Siracusa)</strong><br><strong>- outCT.tar.gz (Catania)</strong></p> <p>####Training Datasets####<br><strong>- t.tar.gz: </strong>Offshore waveforms for gauge sites in a 3D numpy file<br><strong>- dZ.tar.gz: </strong>Local deformation fields in 2D numpy files or reduced 1D flat files on flooded grids<br><strong>- d.tar.gz: </strong>Maximum inundation depth in 2D numpy files or reduced 1D flat files</p> <p>####Flood Mask####<br>- Identifies floodable and non-floodable areas in the domain.<br><strong> - zero_mask_SR_961.npy.gz</strong><br><strong> - zero_mask_CT_892.npy.gz</strong></p> <p>####Flood Indices####<br>Links between 2D and 1D numpy files:<br>- Indices connecting 2D (events x m x n) numpy files to flattened 1D (events x m) numpy files:<br><strong> - zero_indices_SR_961.npy.gz</strong><br><strong> - zero_indices_CT_892.npy.gz</strong><br>- Indices connecting latitude-longitude coordinates to 1D (events x m) numpy files:<br><strong> - lat_lon_idx_CT_892.npy.gz</strong><br><strong> - lat_lon_idx_SR_961.npy.gz</strong></p> <p>The filenames follow nomenclature representing <strong>{parameter}_{site}_{size}</strong></p> <ol> <li>The first var represents the parameter of the file. <ul> <li>d - 2 dimensional file for inundation depth (<strong>events </strong>x <strong>m </strong>x <strong>n</strong>)</li> <li>dZ - 2 dimensional file for local deformation( <strong>events </strong>x <strong>m</strong> x <strong>n</strong>)</li> <li>dflat - 1 dimensional flat file for inundation ( <strong>events </strong>x <strong>locations</strong>)</li> <li>dZflat - 1 dimensional flat file for local deformation( <strong>events </strong>x <strong>locations</strong>)</li> <li>t - 3 dimension offshore waveform for (<strong>events </strong>x <strong>gauges </strong>x <strong>timesteps</strong>)</li> <li>lat_lon_idx - index file for matching lat long coordinate with location indices(<strong>locations x lat x lon</strong>)</li> </ul> </li> <li>The second var represents site of the file. <ul> <li>CT - Catania</li> <li>SR - Siracusa</li> </ul> </li> <li>The third var represents number of events in the file or the selection size.<br> <ul> <li>892 - represents the training events for the file, of 4 different sizes for Catania and Siracusa each</li> </ul> </li> </ol> <div>More information on the attached readme.</div>
Data sets used in "Neural network emulation of the formation of organic aerosols based on the explicit GECKO-A chemistry model"
<p>The training, validation, and testing data sets for toluene, dodecane, and alpha-pinene models described in the manuscript. A link to the manuscript will be added here when it becomes available. All trajectories in the data sets were generated using GECKO-A. The source code for using the data sets can be found at https://github.com/NCAR/gecko-ml </p>
Data archive for paper "Machine Learning Emulation of 3D Cloud Radiative Effects"
<p><strong>Overview</strong></p> <p>This archive contains models, data, and the Singularity image to optionally rerun experiments described in "<a href="https://doi.org/10.1029/2021MS002550">Machine Learning Emulation of 3D Cloud Radiative Effects</a>".</p> <p>For the Python tool to generate synthetic data, please refer to the <a href="https://github.com/dmey/synthia">Synthia repository</a>.</p> <p><strong>Prerequisites</strong></p> <ul> <li>Linux or macOS with Bash shell.</li> <li><a href="https://sylabs.io/singularity/">Singularity</a> (tested with version 3.6.3-1.el8)*.</li> <li><a href="https://en.wikipedia.org/wiki/Portable_Batch_System">Portable Batch System</a> (PBS) job scheduler**.</li> </ul> <p>*Please note that all steps require <a href="https://sylabs.io/">Singularity</a> to be installed on your system. If you are looking for information on how to install or use Singularity, please refer to the <a href="https://sylabs.io/docs">Singularity documentation</a>.</p> <p>**Although PBS in not a strict requirement, it is required to run all helper scripts as included in this repository. Please note that depending on your specific system settings and resource availability, you may need to modify PBS parameters at the top of submit scripts stored in the <code>hpc</code> directory (e.g. <code>#PBS -lwalltime=24:00:00</code>).</p> <p><strong>Initialization</strong></p> <p>Deflate the data archive with:</p> <pre><code>./init.sh </code></pre> <p>Compile ecRad with Singularity:</p> <pre><code>./tools/singularity/compile_ecrad.sh </code></pre> <p><strong>Usage</strong></p> <p>To reproduce the results as described in the paper, run the following commands from the <code>hpc</code> folder:</p> <pre><code>qsub -v JOB_NAME=mlp_default ./submit_grid_search_default.sh qsub -v JOB_NAME=mlp_synthia ./submit_grid_search_synthia.sh qsub submit_benchmark.sh </code></pre> <p>then, to plot stats and identify notebooks run:</p> <pre><code>qsub submit_stats.sh </code></pre> <p><strong>License</strong></p> <p>Paper code released under the <a href="./LICENSE.txt">MIT license</a>. Data released under <a href="./data/LICENSE.txt">CC BY 4.0</a>. <a href="https://confluence.ecmwf.int/display/ECRAD">ecRad</a> released under the <a href="./ecrad/LICENSE">Apache 2.0 license</a>.</p>
Supplementary Data: Sensitivity of Air Pollution Exposure and Disease Burden to Emission Changes in China using Machine Learning Emulation.
<p>Temporary duplicate of:</p> <p>Conibear, L., Reddington, C. L., Silver, B. J., Chen, Y., Arnold, S. R., & Spracklen, D. V. (2022). <em>Supplementary Data: Sensitivity of Air Pollution Exposure and Disease Burden to Emission Changes in China using Machine Learning Emulation. University of Leeds. [Dataset]</em>. https://doi.org/doi.org/10.5518/1055</p>
CLM-FATES parameter estimation using the 'calibrate, emulate, sample' approach - experimental results
Open the record for dataset details and reuse information.
Fast emulator of changes in crop yields at different levels of global warming
<p>This is the Online Supplement to the following publication: Ostberg, S., Schewe, J., Childers, K.,<br> and Frieler, K.: Changes in crop yields and their variability at different levels of global warming, Earth System<br> Dynamics, 9, 2018. The Supplement contains a number of additional figures as well as the emulator coefficients<br> needed to apply the emulators presented in the paper to derive yield changes for any given pair of global mean<br> temperature change (<span class="math-tex">\(\Delta\)</span>GMT) and atmospheric CO<sub>2</sub> concentration (pCO<sub>2</sub>). See readme.pdf for details.</p>
Data and analysis for "Fldgen v1.0: An Emulator with Internal Variability and Space-Time Correlation for Earth System Models"
<p>This is an archive of the raw data and analysis source code for the paper "Fldgen v1.0: An Emulator with Internal Variability and Space-Time Correlation for Earth System Models". The archive contains:</p> <ul> <li><strong>devel.Rmd : </strong>Source code for the worksheet that contains the early development and figures for the paper.</li> <li><strong>devel.html</strong> : HTML rendering of devel.Rmd</li> <li><strong>lg-ensemble-stats.Rmd </strong>: Source code for the worksheet that contains the statistical analysis described in the paper.</li> <li><strong>lg-ensemble-stats.html</strong> : HTML rendering of lg-ensemble-stats.Rmd</li> <li><strong>cc-analysis.Rmd </strong>: Analysis of the compromise conjecture raised by some readers of the paper</li> <li><strong>cc-analysis.nb.html</strong> : HTML rendering of cc-analysis.Rmd</li> <li><strong>data.tar.bz2 </strong>: Input data for the analyses above.</li> </ul> <p>The source code in this archive is written in R and requires the R runtime environment. It also uses the fldgen package, version 1.0.0, which is available at <a href="https://github.com/JGCRI/fldgen">https://github.com/JGCRI/fldgen</a></p> <p> </p>
Data Release: "A neural network emulator of the Advanced LIGO and Advanced Virgo selection function"
<p>This dataset contains results presented in "<strong>A neural network emulator of the Advanced LIGO and Advanced Virgo selection function</strong>" (<a href="https://www.arxiv.org/abs/2408.16828">arXiv: 2408.16828</a>).</p> <p>The code used to generate this data and produce figures in the paper can be found at <a href="https://github.com/tcallister/learning-p-det/">https://github.com/tcallister/learning-p-det/</a>. Specific instructions about the workflow are provided in the <a href="https://tcallister.github.io/learning-p-det/">accompanying documentation</a>.</p> <p>The primary deliverable of this work is a trained neural network emulator for the compact binary selection function during the Advanced LIGO and Advanced Virgo O3 observing run. This emulator is made available in a standalone companion repository, <a href="https://github.com/tcallister/pdet">https://github.com/tcallister/pdet</a>.</p> <p>Additional information:</p> <ul> <li>The files <em>endo3_bbhpop-LIGO-T2100113-v12.hdf5</em>, <em>endo3_bnspop-LIGO-T2100113-v12.hdf5</em>, and <em>endo3_nsbhpop-LIGO-T2100113-v12.hdf5</em>, used for network training, were created and released by the LIGO-Virgo-KAGRA Collaboration at <a href="../records/7890437">https://zenodo.org/records/7890437</a>.</li> <li>The file <em>sampleDict_FAR_1_in_1_yr.pickle</em>, used during hierarchical inference, was created via code in the repository <a href="https://github.com/tcallister/get-lvk-data">https://github.com/tcallister/get-lvk-data</a>.</li> <li>Inference results (<em>popsummary_standardInjections.h5</em> and <em>popsummary_dynamicInjections.h5</em>) are provided in the <em>popsummary</em> results format; see <a href="https://git.ligo.org/christian.adamcewicz/popsummary">https://git.ligo.org/christian.adamcewicz/popsummary</a>.</li> </ul> <p>Changelog:</p> <ul> <li>v2: Added missing file <em>sampleDict_FAR_1_in_1_yr.pickle</em></li> </ul>
Dataset for the plots in paper: 'A Kronecker product accelerated efficient sparse Gaussian Process (E-SGP) for flow emulation' in 'Journal of Computational Physics'
<p>The .xlsx file contains the data used for the plots Fig. 3, 4, 7, 8 and 9 in the paper 'Kronecker product accelerated efficient sparse Gaussian Process (E-SGP) for flow emulation'.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.