Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
136
datasets available to search
ShareScore release 0.7.1
Dataset results
136 results for “diffusion modelling”
Large-eddy simulation investigating the role of double-diffusive convection in basal melting of Antarctic ice shelves: model output
<p>Model output used in the publication:</p> <p>M. G. Rosevear, B. Gayen, B. K. Galton-Fenzi, The role of double-diffusive convection in the basal melting of Antarctic ice shelves. <em>Proc. Natl. Acad. Sci. </em>(2021) https://doi.org/10.1073/pnas.207541118</p> <p>See README.md for a description of the data.</p>
Dataset and neural network weights to the paper: "Generative diffusion for regional surrogate models from sea-ice simulations"
<p>All the needed code and data to reproduce the results from the paper: "Generative diffusion for regional surrogate models from sea-ice simulations".<br>While most of the code is a frozen clone of the original <a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, this capsule also includes the dataset and neural network weights to train and apply the surrogate models.</p> <p>The <strong>dataset</strong> for training and evaluation can be found at <em>data/nextsim</em>, which includes three different Zarr folders for training/validation/testing. The dataset is based on neXtSIM simulation data and ERA5 forcing data and extracted from the <a href="https://ige-meom-opendap.univ-grenoble-alpes.fr/thredds/catalog/meomopendap/extract/catalog.html">SASIP shared data OpenDAP server</a>:</p> <ul> <li>The neXtSIM simulations were performed by Gauillaume Boutin and published in the paper "<a href="https://doi.org/10.5194/tc-17-617-2023">Arctic sea ice mass balance in a new coupled ice–ocean model using a brittle rheology framework</a>" (Boutin et al., 2023) and available as Zenodo <a href="../records/7277523">dataset</a> (Boutin et al., 2022).</li> <li>The forcing data is based on the ERA5 reanalysis dataset published in the paper: "<a href="https://doi.org/10.1002/qj.3803">The ERA5 global reanalysis</a>" (Hersbach et al., 2020) and available as dataset from the Copernicus Climate Change Service (C3S, Copernicus Climate Change Service, 2023). The here used forcing data is based on the <a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels">hourly reanalysis data on single levels</a> and interpolated with nearest neighbors to the curvilinear grid as used in the output from the neXtSIM simulations. <strong>Disclaimer:</strong> The results contain modified Copernicus Climate Change Service information, 2023. Neither the European Commission nor ECMWF is responsible for any use that may be made of the Copernicus information or data it contains.</li> </ul> <p>The <strong>neural network weights</strong> are included under <em>data/models </em>and split into weights for the deterministic models and the diffusion models.<br>These neural network weights have been used to generate the results presented in the paper.</p> <p>In this capsule, the <em>notebooks</em> folder includes also the figures used within the paper and additional trajectory data used in the qualitative analysis of the paper.</p> <p>Generally, we recommend to just download the <em>data.tar.gz </em>file and use otherwise the original <a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, since the here included code can be outdated. We further refer to the repository for additional information.</p> <p> </p> <p>Contained in this capsule:</p> <ul> <li>configs.tar.gz: The configuration files for the experiments.</li> <li>data.tar.gz: The dataset and neural network weights.</li> <li>diffusion_nextsim.tar.gz: The main code for the neural network etc.</li> <li>environment.yaml: The anaconda environment file, can be used to install the needed packages.</li> <li>notebooks.tar.gz: The notebooks that were used to create the figures in the paper. The figures from the paper and the data from the qualitative analysis are included as well.</li> <li>readme.md: The readme file from the repository.</li> <li>scripts.tar.gz: The scripts used for the experiments.</li> <li>setup.py: the file to install the <em>diffusion_nextsim</em> package in a python environment.</li> </ul> <p>References:</p> <p>Guillaume Boutin, Heather Regan, Einar Ólason, Laurent Brodeau, Claude Talandier, Camille Lique, & Pierre Rampal. (2022). Data accompanying the article "Arctic sea ice mass balance in a new coupled ice-ocean model using a brittle rheology framework" (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7277523</p> <p>Boutin, G., Ólason, E., Rampal, P., Regan, H., Lique, C., Talandier, C., Brodeau, L., and Ricker, R.: Arctic sea ice mass balance in a new coupled ice–ocean model using a brittle rheology framework, The Cryosphere, 17, 617–638, https://doi.org/10.5194/tc-17-617-2023, 2023.</p> <p>Copernicus Climate Change Service (2023): ERA5 hourly data on single levels from 1940 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS), DOI: <a href="https://doi.org/10.24381/cds.adbb2d47">10.24381/cds.adbb2d47</a>.</p> <p>Hersbach H, Bell B, Berrisford P, et al. The ERA5 global reanalysis. <em>Q J R Meteorol Soc</em>. 2020; 146: 1999–2049. <a href="https://doi.org/10.1002/qj.3803">https://doi.org/10.1002/qj.3803</a></p> <p> </p>
Data, plotting scripts, and figures for "Assessing diffusion model impacts on enstrophy and flame structure in lean premixed flames"
<p>This repository contains the data, plotting scripts, and figures associated with the paper "Assessing diffusion model impacts on enstrophy and flame structure in lean premixed flames" by Aaron J. Fillo, Peter E. Hamlington, and Kyle E. Niemeyer.</p> <p>See the README file for additional details.</p>
Task 3 Dataset for Dreaming of Electrical Waves: Generative Modeling of Cardiac Excitation Waves using Diffusion Models
Open the record for dataset details and reuse information.
Diffusion coefficients on amorphous polystyrene and modelling of migration levels from plastic packaging
<p>This dataset is actually supplementary data of the scientific article:</p> <p>Martinez-Lopez, Brais; Gontard, Natalie and Peyron, Stephane "Worst case prediction of additives migration from polystyrene for food safety purposes: a model update" in Food Additives and Contaminants Part A, doi:10.1080/19440049.2017.1402129.</p> <p>If you use it, please cite it using the reference file we have provided.</p> <p>This description is the same as in the file "readme.txt", included in the upload.</p> <p>List of files:</p> <ul> <li>The file database_D contains the experimental diffusivity data for amorphous polystyrene used for the figure 1b. It is a spreadsheet file with two tabs. In the first tab, the diffusion coefficients can be found by choosing molecule family (and the publication were they were found) and temperature in celsius degrees. The second tab contains the same diffusivity data, but they are ranged by increasing molecular weight and temperature. This file is available in open document (.ods) and microsoft excel (.xlsx) formats.</li> <li>The file migration modelling is also a spreadsheet file, and contains several tabs. The first tab (diffusion coefficient) is an implementation of equation 1, the predictive model for overestimated diffusion coefficients. The given Ap and tau parameter sets are the ones specified in Table 2 for amorphous polystyrene. The second tab (migration levels) is an implementation of equation 3, the solution to Fick's second law that is used to predict migration levels in food, for pre-selected values of alpha (equation 5). The tabs labeled alpha =... contain the sums used in the equation, whereas the tab "roots" contains the first 200 roots of trascendental equation 4, needed to calculate the sum or terms. This file is also available in open document (.ods) and microsoft excel (.xlsx) formats.</li> <li>The file "table.pdf" sums the main characteristics of the molecule families, together with the references where they were found (in the second page).</li> <li>The file reference.bib contains the reference that should be cited if you use this dataset for your own work.</li> <li>Finally, the file readme.txt contains this very same description.</li> </ul> <p>These files have undergone thorough check, so there should not be any mistakes. In the rare event that you find one, please report it to the author so it can get fixed.</p> <p>bramar@food.dtu.dk</p> <p>Brais Martínez López, PhD<br> Assistant professor<br> DTU Fødevareinstituttet<br> Danmarks Tekniske Universitet<br> Søltofts Plads<br> Bygning 227<br> 2800 Kgs. Lyngby</p> <p> </p> <p> </p> <p> </p>
Science ready spectra and their best-fitting models described in the research paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al.
<p>Science ready spectra of nine ultra-diffuse galaxies in the Coma cluster collected with the Binospec multi-object spectrograph and their best-fitting PEGASE.HR templates obtained using the NBursts full spectrum fitting code. These spectra were presented in the paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al. accepted for publication in the Astrophysical Journal on Sep/3/2019 (arXiv:1901.05489).</p> <p>Each spectrum is presented as a binary FITS table, which contains a spectrum (wavelength, flux, uncertainties), best-fitting template, best-fitting parameters (radial velocity, age, metallicity), and a pixel mask used in the fitting procedure. For six galaxies there are two files provided: (i) one-dimensional optimally extracted integrated spectrum and (ii) two dimensional spectrum for spatially resolved radial velocity information. For the remaining three galaxies, only spatially resolved spectra are provided.</p>
Dataset and structure database for an ML model to predict diffusivity in ZIF variants
<p>This dataset accompanies the publication titled "Data Mining for Predicting Gas Diffusivity in Zeolitic-imidazolate Frameworks (ZIFs)" (DOI: <a href="https://doi.org/10.1039/D2TA02624D">https://doi.org/10.1039/D2TA02624D</a>)</p> <p><a href="https://zenodo.org/api/files/b80f6d07-3bf4-484c-97ac-5d579fb0cc27/ESI_2_dataset.xlsx?versionId=dc4525d0-1c5c-478c-9bef-a156587ad69b">ESI_2_dataset.xlsx</a>: Descriptors for all ZIFs of the publication and simulations output, in the form of diffusivities of gas molecules (He up to iso-butane), in all ZIFs.</p> <p>ZIF_database.zip: ZIP file containing all ZIFs prepared by the authors (as discussed in the publication), through various units replacements, in the SOD topology, in .pdb format.</p>
Data and code for gmd-2023-113 "Parameter estimation for ocean background vertical diffusivity coefficients in the Community Earth System Model (v1.2.1) and its impact on ENSO forecast"
<p>Data and code for the paper "Parameter estimation for ocean background vertical diffusivity coefficients in the Community Earth System Model (v1.2.1) and its impact on ENSO forecast"</p> <p>includes: </p> <p>The model is Community Earth System Model (v1.2.1) (provided by www.cesm.ucar.edu)</p> <p>Data assimilation code is initially provided by Data Assimilation Research Testbed (DART) (https://dart.ucar.edu/), some modifications are made to enable parameter estimation function of ocean background vertical diffusivity coefficients. And the programs and scripts for deal with OISST and EN4 profiles are also developed.</p> <p>The parameter sensitivity experiment results are saved as <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/sensitive2008-2012.nc">sensitive2008-2012.nc</a> and <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/sensitive2008-2012salt.nc">sensitive2008-2012salt.nc</a> for temperature and salinity, respectively. And the python script to draw the results is </p> <p>The state estimation results are provided as <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/Temp_05-17.nc">Temp_05-17.nc</a> and <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/Temp_05-17.nc">Salt_05-17.nc</a> for temperature and salinity, respectively.</p> <p>The parameter estimation results are provided as <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/PE_Temp_05-17.nc">PE_Temp_05-17.nc</a> and <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/PE_Temp_05-17.nc">PE_Salt_05-17.nc</a> for temperature and salinity, respectively.</p> <p>the estimated paremeter ensemble is saved in <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/parameters.nc">parameters.nc</a></p> <p>the python script for comparing the SE and PE results is <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/plot_analysis.py">plot_analysis.py</a></p> <p>the nino3.4 indices computed by the forecast experiment is saved in <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/fcst_correlation.nc">fcst_correlation.nc</a></p> <p> </p>
Datasets and trained diffusion models for "Diffusion Models for Interferometric Satellite Aperture Radar"
<p>A set of trained Probabilistic Diffusion Models (PDMs) and corresponding training datasets for the paper "<a href="https://doi.org/10.48550/arXiv.2308.16847">Diffusion Models for Interferometric Satellite Aperture Radar</a>", by Tuel, Kerdreux et al. The code for this paper can be found at <a href="https://github.com/thomaskerdreux/PDM_SAR_InSAR_generation">this link</a>.</p> <p><strong>Training datasets</strong></p> <p>- "InSAR_noise_32x32.zip": a dataset of 32x32 ground deformation scenes obtained from InSAR interferograms over New Mexico with the small baseline subset (SBAS) algorithm. Images were normalised to [0, 1].</p> <p>- "insar_unwrapped_phase_normalised.zip": a dataset of 128x128 InSAR interferograms obtained from Sentinel-1 acquisitions over Nex Mexico. Images were normalised to [0, 1].</p> <p><strong>Trained Models</strong></p> <p>We provide 6 trained PDMs in separate .zip files. Each .zip file contains the model weights (in *.pt format) and the model metadata file (in *.json format).</p> <p>- "mnist_32_cond_sigma_100.zip": a class-conditional model trained with 100 diffusion time steps on 32x32 MNIST images;</p> <p>- "mnist_32_no_cond_sigma_100.zip": an unconditional model trained with 100 diffusion time steps on 32x32 MNIST images;</p> <p>- "SAR_lowres_128_cond_sigma_2000.zip": a low-resolution (256 to 128) model trained with 2000 diffusion time steps on 128x128 TenGeoP-SARwv images;</p> <p>- "SAR_superres_128_to_256_cond_sigma_2000.zip": a super-resolution (128 to 256) model trained with 2000 diffusion time steps on TenGeoP-SARwv images;</p> <p>- "insar_phase_128_sigma_2000.zip": an unconditional model trained with 2000 diffusion time steps on 128x128 Sentinel-1 InSAR interferograms over New Mexico;</p> <p>- "insar_noise_32_sigma_1000.zip": an unconditional model trained with 1000 diffusion time steps on 32x32 Sentinel-1 InSAR ground deformation scenes over New Mexico.</p>
Scripts, models, and data for manuscript "On the Role of Stern- and Diffuse-Layer Polarization Mechanisms in Porous Media"
<p>This repository contains Matlab scripts, Comsol Multiphysics models, and numerical simulation data used to generate the plots in the manuscript</p> <p>Bücker, M., Flores Orozco, A., Undorf, S., and Kemna, A., 2019, <em>On the Role of Stern- and Diffuse-Layer Polarization Mechanisms in Porous Media</em>, submitted to JGR: Solid Earth.</p> <p>If you find this data useful in your own research, please cite this manuscript.</p>
Summary of measured and modeled light curve parameters for diffuse, direct, and intermediate light curves for 14 whole-canopy 1mx1m plots sampled near the shrub LTER sites at Toolik Field Station, Alaska, summer 2012.
14 1m x 1m shrub plots were sampled the summer of 2012 under direct and diffuse light conditions. Light response curves were measured under each light condition for each plot using a Li-Cor 6400 to measure net ecosystem exchange (NEP); these measurements were modelled using a saturatingMichaelis-Menton formula. The best fit parameters for those models are contained here (Pmax, K, RE, Eo, and light compensation point) for each individual NEP light response curve (direct, diffuse, and intermediate light conditions) measured with corresponding NDVI , LAI, diffuse light fraction, and average temperature. Sorting variables and curve ID numbers for each curve match the corresponding data in the flux data file.
Examples of ALD saturation profiles in rectangular channel LHAR structures simulated with a diffusion-reaction model
<p>Examples of atomic layer deposition ALD saturation profiles in rectangular channel LHAR structures simulated with a diffusion-reaction model by Ylilammi et al. (Journal of Applied Physics <strong>123</strong>, 205301 (2018); <a href="https://doi.org/10.1063/1.5028178">https://doi.org/10.1063/1.5028178</a>) as function of (a) Reactant A pulse time <em>t</em>, (b) Reactant A partial pressure <em>p</em>, and (c) sticking coefficient <em>c</em>. Parameters used in the simulation, if not otherwise stated: channel height <em>H</em> 500 nm, temperature 250°C, 250 cycles, inert carrier gas partial pressure 500 Pa, Reactant A molar mass 100 g/mol, inert carrier gas molar mass 28 g/mol, Reactant A diameter 0.600 nm, inert gas diameter 0.374 nm, adsorption capacity 4 metal atoms per nm<sup>2</sup>, density of the material grown 3.5 g/cm<sup>3</sup>, pulse time 0.1 s, Reactant A partial pressure 100 Pa, sticking coefficient 0.01. </p> <p>Abbreviations: ALD = atomic layer deposition, LHAR = lateral high aspect ratio</p>
Synthbuster: Towards Detection of Diffusion Model Generated Images
<p>Dataset described in the paper "Synthbuster: Towards Detection of Diffusion Model Generated Images" (Quentin Bammey, 2023, <i>Open Journal of Signal Processing</i>)</p><p>This dataset contains synthetic, AI-generated images from 9 different models:</p><ul><li>DALL·E 2</li><li>DALL·E 3</li><li>Adobe Firefly</li><li>Midjourney v5</li><li>Stable Diffusion 1.3</li><li>Stable Diffusion 1.4</li><li>Stable Diffusion 2</li><li>Stable Diffusion XL</li><li>Glide</li></ul><p> </p><p>1000 images were generated per model. The images are loosely based on raise-1k images (Dang-Nguyen, Duc-Tien, et al. "Raise: A raw images dataset for digital image forensics." Proceedings of the 6th ACM multimedia systems conference. 2015.). For each image of the raise-1k dataset, a description was generated using the Midjourney /describe function and CLIP interrogator (https://github.com/pharmapsychotic/clip-interrogator/). Each of these prompts was manually edited to produce results as photorealistic as possible and remove living persons and artists names.</p><p> </p><p>In addition to this, parameters were randomly selected within reasonable values for methods requiring so.</p><p>The prompts and parameters used for each method can be found in the `prompts.csv` file.</p><p> </p><p>This dataset can be used to evaluate AI-generated image detection methods. We recommend matching the generated images with the real Raise-1k images, to evaluate whether the methods can distinguish the two of them. Raise-1k images are not included in the dataset, they can be downloaded separately at (http://loki.disi.unitn.it/RAISE/download.html).</p><p> </p><p>None of the images suffered degradations such as JPEG compression or resampling, which leaves room to add your own degradations to test robustness to various transformation in a controlled manner.</p><p> </p>
Synthetic IAM with latent diffusion models
<p>SyntheticHTR: Handwritten Text Image Synthesis based on Latent Diffusion Models</p>
Sensitivity of a Coarse-Resolution Global Ocean Model to a Spatially Variable Neutral Diffusivity - ACCESS-OM2 data and plotting routines
<p>This repository contains the processed data and plotting routines associated with the article</p> <p>Holmes, Groeskamp, Stewart and McDougall (2022), Sensitivity of a Coarse-Resolution Global Ocean Model to a Spatially Variable Neutral Diffusivity, Journal of Advances in Modeling Earth Systems (JAMES), doi: 10.1029/2021MS002914, http://dx.doi.org/10.1029/2021MS002914</p> <p>The contents includes post-processed data output from the 1-degree ACCESS-OM2 ocean-sea-ice model simulations and the python/jupyter plotting routines required to make the plots.</p> <p>The processing script is Holmes2022JAMES_Neutral_Diffusion_ACCESS-OM2_Plotting_Script.ipynb. The data files consist of time-averages or time series of certain metrics processed using NCO tools from the raw ACCESS-OM2 simulation output.</p>
Deep Learning for Reaction-Diffusion Glioma Growth Modeling: Towards a Fully Personalized Model? — Supporting Data
<p>Supporting data for Martens et al. Deep Learning for Reaction-Diffusion Glioma Growth Modelling: Towards a Fully Personalised Model? arXiv:2111.13404.</p>
DiffModeler: Large Macromolecular Structure Modeling in Low-Resolution Cryo-EM Maps Using Diffusion Model
<p>Here, we store the modeled structures generated by DiffModeler for its 4 benchmark datasets: CryoREAD dataset(0-5A resolution, protein-DNA/RNA complex), ModelAngelo dataset(0-5A resolution, most protein complexes, a few protein-RNA complex), intermediate resolution dataset (5-10A resolution, protein complex), low resolution dataset (10-20A resolution, protein complex). For all protein-DNA/RNA complex, the map will be modeled by CryoREAD+DiffModeler.</p> <p>For each dataset, we keep the modeled structures by DiffModeler, named as [EMD-ID]_DiffModeler.cif; and their corressponding native structures from RCSB are saved as [EMD-ID]_[PDB_ID]_native.cif.</p> <p>For CryoREAD dataset, it includes 61 targets. For ModelAngelo dataset, it includes 28 targets. For intermediate resolution dataset , it includes 71 targets. For low resolution dataset, it incldues 6 targets.</p> <p>For intermediate resolution dataset, many maps were run with inaccurate AF2 predicted single-chain structures. We also benchmarked DiffModeler's performance by using native single-chain structures as input. They are saved under "dataset_5_10A_nativechain" folder.</p> <p>Additionally, we have stored the traced backbone map of the intermediate resolution dataset in the "dataset_5_10A_diffusion_traced_backbone_map" folder. The traced maps are saved as [EMD-ID]_diffusion.mrc. The intermediate reverse diffusion maps of the intermediate resolution dataset are saved in the "dataset_5_10A_reverse_diffusion_maps" folder. Each sub-folder is named according to the corresponding map's [EMD-ID] and contains three intermediate reverse diffusion maps: 20percentile_reverse_diffusion.mrc, 50percentile_reverse_diffusion.mrc, and 80percentile_reverse_diffusion.mrc. A higher percentile indicates a map closer to the end of the reverse diffusion steps.</p> <p>If you used DiffModeler, please cite: "Wang, Xiao, Han Zhu, Genki Terashi, Manav Taluja, and Daisuke Kihara. "DiffModeler: Large Macromolecular Structure Modeling in Low-Resolution Cryo-EM Maps Using Diffusion Model." bioRxiv (2024): 2024-01.".</p> <p>If you used CryoREAD, please cite: "Xiao Wang, Genki Terashi & Daisuke Kihara. De novo structure modeling for nucleic acids in cryo-EM maps using deep learning. Nature Methods, 2023."</p>
Surrogate Modeling Benchmark - Two-dimensional heat diffusion model
<p>This dataset is related to the Two-dimensional heat diffusion model benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld: <a href="https://uqworld.org/t/benchmark-case-two-dimensional-heat-diffusion-model/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-two-dimensional-heat-diffusion-model/</a>.</p> <p>The experimental designs include datasets with 400, 800, 1200, 1600, and 2000 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of <em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Matérn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Matérn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual – Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual – Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual – Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em> (the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Zürich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Zürich.</p>
Surrogate Modeling Benchmark - One-dimensional diffusion model
<p>This dataset is related to the One-dimensional diffusion model benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld: <a href="https://uqworld.org/t/benchmark-case-one-dimensional-diffusion-model" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-one-dimensional-diffusion-model</a>.</p> <p>The experimental designs include datasets with 200, 400, 600, 800, and 1000 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of <em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Matérn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Matérn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual – Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual – Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual – Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em> (the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Zürich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Zürich.</p>
Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms Dataset
<p>This is the official data repository for the paper: "Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms", available at:<a href="https://www.arxiv.org/abs/2409.19371"> https://www.arxiv.org/abs/2409.19371</a>. The corresponding code is available at: <a href="https://github.com/david-stojanovski/EDMLX">https://github.com/david-stojanovski/echo_from_noise</a></p> <p> </p> <p>The synthetic data is produced using a variety of generative architectures, including the <strong>Elucidating Diffusion Model (EDM), Variance Exploding (VE), Variance Preserving (VP)</strong>, and our novel models, <strong>EDM-L64</strong> and <strong>EDM-L128</strong>, which employ <strong>latent diffusion</strong> strategies to significantly reduce computational cost. By incorporating <strong>spatially adaptive normalization (SPADE) blocks</strong> and <strong>Γ-distribution-based Variational Autoencoders (Γ-VAE)</strong>, these datasets ensure that the generated images preserve the essential semantic features required for training deep learning models.</p> <p> </p> <p>All pretrained classification and segmentation models can be found within the <strong>trained_models </strong>file.</p> <p>All generated images can be found within the <strong>generated_data </strong>file. Included is the <strong>CAMUS</strong> and original <strong>Semantic Diffusion Model (SDM) </strong>data, as well as a folder labelled <strong>easy_inference</strong> designed to contain all relevant labelmaps in a convenient folder for generating replicas of the dataset (detailed at codebase).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.