Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

59

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

59 results for “Surrogate model”

Learn how ShareScore rates datasets ↗
zenodo52/100

Dataset for training the Surrogate Model of microlaser neurons on the reduced MNIST classification task

<p>This dataset was used to train a surrogate multilayer perceptron surrogate model of microlaser neurons.</p> <p>It is in csv format. It was generated using the Yamada Model as found in&nbsp;</p> <p><span>Selmi F, Braive R, Beaudoin G, Sagnes I, Kuszelewicz R and Barbay S 2014 Relative Refractory Period in an Excitable Semiconductor Laser <em>Phys. Rev. Lett.</em> <strong>112</strong> 183902</span>.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Datasets from study: "Land surface observations boost temperature forecast skill: experiments using Long Short-Term Memory surrogate for physics-based models to assess potential predictability"

<p>This repository contains the datasets needed to reproduce the figures from manuscript: Land surface observations boost temperature forecast skill: experiments using Long Short-Term Memory surrogate for physics-based models to&nbsp;assess potential predictability.</p> <p>In this study, we examine the potential of land surface temperature and vegetation data, which are not routinely assimilated in NWP models, for enhancing temperature forecast skill. We build surrogate models for NWP using Long Short-Term Memory.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Dataset and neural network weights to the paper: "Generative diffusion for regional surrogate models from sea-ice simulations"

<p>All the needed code and data to reproduce the results from the paper: "Generative diffusion for regional surrogate models from sea-ice simulations".<br>While most of the code is a frozen clone of the original&nbsp;<a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, this capsule also includes the dataset and neural network weights to train and apply the surrogate models.</p> <p>The <strong>dataset</strong> for training and evaluation can be found at&nbsp;<em>data/nextsim</em>, which includes three different Zarr folders for training/validation/testing. The dataset is based on neXtSIM simulation data and ERA5 forcing data and extracted from the <a href="https://ige-meom-opendap.univ-grenoble-alpes.fr/thredds/catalog/meomopendap/extract/catalog.html">SASIP shared data OpenDAP server</a>:</p> <ul> <li>The neXtSIM simulations were performed by Gauillaume Boutin and published in the paper "<a href="https://doi.org/10.5194/tc-17-617-2023">Arctic sea ice mass balance in a new coupled ice&ndash;ocean model using a brittle rheology framework</a>" (Boutin et al., 2023) and available as Zenodo <a href="../records/7277523">dataset</a> (Boutin et al., 2022).</li> <li>The forcing data is based on the ERA5 reanalysis dataset published in the paper: "<a href="https://doi.org/10.1002/qj.3803">The ERA5 global reanalysis</a>" (Hersbach et al., 2020) and available as dataset from the Copernicus Climate Change Service (C3S, Copernicus Climate Change Service, 2023). The here used forcing data is based on the <a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels">hourly reanalysis data on single levels</a> and interpolated with nearest neighbors to the curvilinear grid as used in the output from the neXtSIM simulations. <strong>Disclaimer:</strong> The results contain modified Copernicus Climate Change Service information, 2023. Neither the European Commission nor ECMWF is responsible for any use that may be made of the Copernicus information or data it contains.</li> </ul> <p>The <strong>neural network weights</strong> are included under <em>data/models </em>and split into weights for the deterministic models and the diffusion models.<br>These neural network weights have been used to generate the results presented in the paper.</p> <p>In this capsule, the <em>notebooks</em> folder includes also the figures used within the paper and additional trajectory data used in the qualitative analysis of the paper.</p> <p>Generally, we recommend to just download the <em>data.tar.gz </em>file and use otherwise the original <a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, since the here included code can be outdated. We further refer to the repository for additional information.</p> <p>&nbsp;</p> <p>Contained in this capsule:</p> <ul> <li>configs.tar.gz: The configuration files for the experiments.</li> <li>data.tar.gz: The dataset and neural network weights.</li> <li>diffusion_nextsim.tar.gz: The main code for the neural network etc.</li> <li>environment.yaml: The anaconda environment file, can be used to install the needed packages.</li> <li>notebooks.tar.gz: The notebooks that were used to create the figures in the paper. The figures from the paper and the data from the qualitative analysis are included as well.</li> <li>readme.md: The readme file from the repository.</li> <li>scripts.tar.gz: The scripts used for the experiments.</li> <li>setup.py: the file to install the <em>diffusion_nextsim</em> package in a python environment.</li> </ul> <p>References:</p> <p>Guillaume Boutin, Heather Regan, Einar &Oacute;lason, Laurent Brodeau, Claude Talandier, Camille Lique, &amp; Pierre Rampal. (2022). Data accompanying the article "Arctic sea ice mass balance in a new coupled ice-ocean model using a brittle rheology framework" (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7277523</p> <p>Boutin, G., &Oacute;lason, E., Rampal, P., Regan, H., Lique, C., Talandier, C., Brodeau, L., and Ricker, R.: Arctic sea ice mass balance in a new coupled ice&ndash;ocean model using a brittle rheology framework, The Cryosphere, 17, 617&ndash;638, https://doi.org/10.5194/tc-17-617-2023, 2023.</p> <p>Copernicus Climate Change Service (2023): ERA5 hourly data on single levels from 1940 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS), DOI:&nbsp;<a href="https://doi.org/10.24381/cds.adbb2d47">10.24381/cds.adbb2d47</a>.</p> <p>Hersbach H, Bell B, Berrisford P, et al. The ERA5 global reanalysis. <em>Q J R Meteorol Soc</em>. 2020; 146: 1999&ndash;2049. <a href="https://doi.org/10.1002/qj.3803">https://doi.org/10.1002/qj.3803</a></p> <p>&nbsp;</p>

openmit-licenseApr 2024View details →
zenodo44/100

Surrogate-based optimization using an artificial neural network for a parameter identification in a 3D marine ecosystem model

<p><strong>Abstract:</strong></p> <p>Parameter identification for marine ecosystem models is important for the assessment and validation of marine ecosystem models against observational data. The surrogate-based optimization (SBO) is a computationally efficient method to optimize complex models. SBO replaces the computationally expensive (high-fidelity) model by a surrogate constructed from a less accurate but computationally cheaper (low-fidelity) model in combination with an appropriate correction approach, which improves the accuracy of the low-fidelity model. To construct a computationally cheap low-fidelity model, we tested three different approaches to compute an approximation of the annually periodic solution (i.e., a steady annual cycle) of a marine ecosystem model: firstly, a reduced number of spin-up iterations (several decades instead of millennia), secondly, an artificial neural network (ANN) approximating the steady annual cycle and, finally, a combination of the both approaches. Except for the low-fidelity model using only the ANN, the SBO yielded a solution close to the target and reduced the computational effort significantly. If an ANN approximating appropriately a marine ecosystem model is available, the SBO using this ANN as low-fidelity model presents a promising and computational efficient method for the validation.</p> <p>&nbsp;</p> <p><strong>Content:</strong></p> <ul> <li>SQLite database including the data of the different optimization runs</li> <li>Structure and weights of the used artificial neural network</li> <li>Tracer concentrations obtain from the high-fidelity model for the different optimization runs</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Phononic crystals dataset for supervised training of surrogate deep learning model

<p>The dataset contains shapes of unit cells of phononic crystals (inputs) in the form of images and corresponding dispersion diagrams (outputs). The dataset is used for deep learning (DL) model training.<br> Outputs are in the form of .mat files which contain vectors of reduced wavevector and corresponding frequencies, and also displacements u, v, w which can be used for polarization calculation.</p> <p>The dataset contains 11000 cases.</p> <p>Note: Ignore names &quot;labels&quot; as these are actually inputs to the DL model, not labels.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Supplementary material for the publication: "Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties"

<p><span><span><span>This dataset contains supplementary code, images and models for the publication &bdquo;Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties&ldquo;.</span></span></span></p> <p>&nbsp;</p> <p><span><span><span>The content will be updated and additionally linked to the corresponding git repositories.</span></span></span></p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

CoUDlabs_WP8_T812_EAWAG_001. Sediment depth measurements for surrogate modeling of sediment build-up in gully pots using temperature data

<p>This dataset contains the results of the&nbsp;experimental campaign and how data were collected on the the <a href="https://co-udlabs.eu/">Co-UDlabs</a> <strong>Work Package 8 (Joint Research Activity 3)</strong>: <i>Improving resilience and sustainability in urban drainage solutions</i>; <strong>Task 8.1</strong>: <i>Development of consensus on measurement of hydraulic and water quality performance of urban drainage technologie</i>s; <strong>Subtask 8.1.2</strong>: <i>Development of scalable measurement protocols to assess the pollutant retention and release potential of urban drainage structures</i>.&nbsp;</p><p>Co-UDlabs is a project funded by the European Union's Horizon 2020 research and innovation programme under grant agreement No 101008626.</p><p>This database was developed as part of the Master Thesis in Environmental Engineering at ETH Zurich (Switzerland). Fuchs, L. (2023). Automated surrogate model to estimate sediment accumulation from temperatures in urban drainage systems. MSc Thesis, ETH Zurich. https://polybox.ethz.ch/index.php/s/IyiM38rRy1vlHWD. Accessed on 10th of October of 2023.</p>

opencc-by-nc-4.0Nov 2023View details →
zenodo40/100

Dataset for Surrogate Model Benchmarking for Dynamic Climate Impact Models

<p>The data represents time series of seasonal weather forecasts for rainfall and temperature. The dataset contains 10 forecasts of 6-month horizon from, two per year, from 2017 to 2021; start dates January 1 and July 1, respectively. Each forecast comprises 50 ensemble members. In total, this sums up to 91300 data points, each containing daily average rainfall, temperature.</p> <p>Each sample (row) comprises following features (columns):</p> <ul> <li><strong>datetime</strong>: Date of the forecast sample.</li> <li><strong>forecast</strong>: Identifier of the ensemble member, i.e. integer between 1 and total number of ensemblemembers.</li> <li><strong>precip</strong>: Averaged daily rainfall forecast in millimeters.</li> <li><strong>temp</strong>: Averaged daily temperature forecast in degree Celsius.</li> </ul> <p>Dataset created by The Weather Company, an IBM business. This service is based on data and products of the European Center for Medium-range Weather Forecasts (ECMWF-Archive and ECMWF-RT). Generated using Copernicus Climate Change Service information [2019 and ongoing]. ECMWF Archive data published under a Creative Commons Attribution 4.0 International (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/<br> Disclaimer: Neither the European Commission nor ECMWF is responsible for any use that may be made of the information it contains.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

A physically interpretable data-driven surrogate model for wake steering

<p>PALM input files for the simulations performed in the study &quot;A physically interpretable data-driven surrogate model for wake steering&quot;&nbsp; by Sengers et al. (2022).&nbsp;</p> <p>The PALM code is available at&nbsp;<a href="https://palm.muk.uni-hannover.de/">https://palm.muk.uni-hannover.de</a><br> Additional information to the input files is given in the README file</p> <p>Cite this as:<br> B.A.M. Sengers (2022). Dataset:&nbsp;A physically interpretable data-driven surrogate model for wake steering. https://doi.org/10.5281/zenodo.6821164</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Surrogate waveform model data for black hole binary systems computed in point-particle black hole perturbation theory

<p>This repository contains all publicly available surrogate data for gravitational waveforms produced within the point-particle black hole perturbation theory framework and calibrated to numerical relativity simulations performed with the Spectral Einstein Code (SpEC).&nbsp;</p> <p>Several surrogate models are currently available in this catalog:</p> <ol> <li><strong>BHPTNRSur2dq1e3</strong>, for aligned spin black hole binary systems with mass-ratios varying from 3 to 1000 and spins from &minus;0.8&le;&chi;1&le;0.8 on the larger black hole. This surrogate model is trained on waveform data generated by point-particle black hole perturbation theory (ppBHPT) with calibration to numerical relativity (NR) data. The waveforms include all spin-weighted spherical harmonic modes up to&nbsp;ℓ=4&nbsp;except the&nbsp;(4,1)&nbsp;and&nbsp;m=0 modes. Model details can be found in <a href="https://arxiv.org/abs/2407.18319">Rink et al. 2024</a>. This data file is used to evaluate the surrogate model with either stand-alone Python code hosted by the <a href="https://bhptoolkit.org/BHPTNRSurrogate/">Black Hole Perturbation Toolkit</a> (Jupyter notebook <a href="https://github.com/BlackHolePerturbationToolkit/BHPTNRSurrogate/blob/main/tutorials/BHPTNRSur2dq1e3.ipynb">tutorial</a>) or the GWSurrogate Python package, which can be found on <a href="https://pypi.python.org/pypi/gwsurrogate/">PyPI</a>&nbsp;or <a href="https://anaconda.org/conda-forge/gwsurrogate">conda-forge</a>.</li> <li><strong>BHPTNRSur1dq1e4</strong>, an updated version of the&nbsp;<strong>EMRISur1dq1e4&nbsp;</strong>model described below. The updated version includes better calibration to NR, a smoother transition to plunge model, and more harmonic modes.&nbsp;Model details can be found in <a href="https://arxiv.org/abs/2204.01972">Islam&nbsp;et al. 2022</a>. This data file is used to evaluate the surrogate model with either stand-alone Python code hosted by the <a href="https://bhptoolkit.org/BHPTNRSurrogate/">Black Hole Perturbation Toolkit</a> (Jupyter notebook <a href="https://github.com/BlackHolePerturbationToolkit/BHPTNRSurrogate/tree/main/tutorials/BHPTNRSur1dq1e4">tutorial</a>) or the GWSurrogate Python package, which can be found on <a href="https://pypi.python.org/pypi/gwsurrogate/">PyPI</a>&nbsp;or <a href="https://anaconda.org/conda-forge/gwsurrogate">conda-forge</a>.</li> <li><strong>EMRISur1dq1e4</strong>,&nbsp;for non-spinning black hole binary systems with mass-ratios varying from 3 to 10000. This surrogate model is trained on waveform data generated by point-particle black hole perturbation theory (ppBHPT), with the total mass rescaling parameter tuned to NR simulations.&nbsp;Available modes are [(2,2), (2,1), (3,3), (3,2), (3,1), (4,4), (4,3),&nbsp;(4,2), (5,5), (5,4), (5,3)]. The m&lt;0 modes are deduced from the m&gt;0 modes. Model details can be found in <a href="https://arxiv.org/abs/1910.10473">Rifat et al. 2019</a>. This data file&nbsp;is used to evaluate&nbsp;the surrogate model with either stand-alone Python code hosted by the <a href="http://github.com/BlackHolePerturbationToolkit/EMRISurrogate">Black Hole Perturbation Toolkit</a>&nbsp;(Jupyter notebook&nbsp;<a href="https://github.com/BlackHolePerturbationToolkit/EMRISurrogate/blob/master/EMRISur1dq1e4.ipynb">tutorial</a>) or the GWSurrogate Python package (Jupyter notebook <a href="https://github.com/sxs-collaboration/gwsurrogate/blob/master/tutorial/notebooks/nonspinning_nr_emri.ipynb">tutorial</a>), which can be found on&nbsp;<a href="https://pypi.python.org/pypi/gwsurrogate/">PyPI</a>.</li> </ol>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Ishigami function

<h1>Surrogate Modeling Benchmark - Ishigami function</h1> <p>This dataset is related to the Ishigami function benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld: <a href="https://uqworld.org/t/benchmark-case-ishigami-function/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-ishigami-function/</a>.</p> <p>The experimental designs include datasets with 40, 80, 120, 160, and 200 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Two-dimensional heat diffusion model

<p>This dataset is related to the Two-dimensional heat diffusion model benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-two-dimensional-heat-diffusion-model/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-two-dimensional-heat-diffusion-model/</a>.</p> <p>The experimental designs include datasets with 400, 800, 1200, 1600, and 2000 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - 100D function

<p>This dataset is related to the 100D function benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-100d-function/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-100d-function/</a>.</p> <p>The experimental designs include datasets with 400, 800, 1200, 1600, and 2000 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Morris function

<p>This dataset is related to the Morris function benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-morris-function" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-morris-function</a>.</p> <p>The experimental designs include datasets with 400, 800, 1200, 1600, and 2000 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - One-dimensional diffusion model

<p>This dataset is related to the One-dimensional diffusion model benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-one-dimensional-diffusion-model" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-one-dimensional-diffusion-model</a>.</p> <p>The experimental designs include datasets with 200, 400, 600, 800, and 1000 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Damped Oscillator

<p>This dataset is related to the damped oscillator benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-damped-oscillator/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-damped-oscillator/</a>.</p> <p>The experimental designs include datasets with 400, 800, 1200, 1600, and 2000 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Truss model

<p>This dataset is related to the truss model benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-truss-model/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-truss-model/</a>.</p> <p>The experimental designs include datasets with 100, 200, 300, 400, and 500 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Wing weight function

<p>This dataset is related to the wing weight function benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-wing-weight-function/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-wing-weight-function/</a>.</p> <p>The experimental designs include datasets with 100, 200, 300, 400, and 500 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Borehole function

<p>This dataset is related to the Borehole function benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld:&nbsp;<a href="https://uqworld.org/t/benchmark-case-borehole-function/">https://uqworld.org/t/benchmark-case-borehole-function/</a>.</p> <p>The experimental designs include datasets with 40, 80, 120, 160, and 200 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Surrogate Modeling Benchmark - Undamped Oscillator

<p>This dataset is related to the undamped oscillator benchmark case. A detailed description of the benchmark case can be found on the public online community website UQWorld: <a href="https://uqworld.org/t/benchmark-case-undamped-oscillator/" target="_blank" rel="noopener">https://uqworld.org/t/benchmark-case-undamped-oscillator/</a>.</p> <p>The experimental designs include datasets with 40, 80, 120, 160, and 200 samples, each generated using optimized maximin distance Latin Hypercube Sampling (LHS) with 1000 iterations. Each dataset is replicated 20 times. The validation set contains 100,000 samples generated by Monte Carlo simulation. Each dataset contains input samples and the corresponding computational model responses.</p> <h2>Description of the dataset file</h2> <p>The dataset file includes two variables:</p> <ul> <li><em>ExpDesigns</em>, and</li> <li><em>ValidationSet</em>.</li> </ul> <p>Both variables are Matlab structures with fields <em>X</em>, <em>Y</em>, and <em>nSamples</em>. Variable <em>ExpDesigns</em> is a non-scalar structure sized according to the number of experimental design groups. Each field of&nbsp;<em>X</em> for the i-th element of the struct array contains replicated datasets, forming a matrix of size [number of samples] x [dimensionality] x [number of replications]. Similarly, each field of Y for the i-th element contains replicated computational model responses that correspond to the experimental design of the same replication, sized [number of samples] x [number of model outputs] x [number of replications]. The same structure logic applies to the <em>ValidationSet</em> variable, except it contains only one dataset per benchmark case.</p> <p>The structure can be summarized as follows:</p> <ul> <li>ExpDesigns(i).X(j,k,l) <ul> <li>i: dataset group,</li> <li>j: sample index,</li> <li>k: variable index, and</li> <li>l: replication index.</li> </ul> </li> </ul> <ul> <li>ExpDesigns(i).Y(j,m,l) <ul> <li>i, j, l: same as above,</li> <li>m: computational model output index.</li> </ul> </li> </ul> <ul> <li>ValidationSet.X(j,k) <ul> <li>j, k: same as above.</li> </ul> </li> </ul> <ul> <li>ValidationSet.Y(j,m) <ul> <li>j, m: same as above.</li> </ul> </li> </ul> <h2>Description of benchmarked metamodel competitors</h2> <p>The selection of competitors was based on our experience with meta-modeling and includes various metamodel types: Polynomial Chaos Expansions (PCE), Polynomial Chaos Kriging (PCK), and Kriging. Given that each metamodel has many hyperparameters, we chose the most general settings to address different benchmark case difficulties, including dimensionality, nonlinearity, and non-monotonicity.</p> <p>For <strong>Polynomial Chaos Expansions (PCE)</strong>, we used a polynomial degree and q-norm adaptivity approach. This approach adaptively increases the maximum polynomial degree and truncation q-norm until the estimated leave-one-out error starts increasing. Maximum polynomial interaction terms were limited to 2 due to the memory requirements for large model dimensionality and large experimental designs. We tested three different solvers to calculate the PCE coefficients: Least Angle Regression (LARS), Orthogonal Matching Pursuit (OMP), and Subspace Pursuit (SP).</p> <p><strong>Polynomial Chaos Kriging (PCK)</strong> employs a sequential combination strategy of PCE and Kriging. PCE uses degree adaptivity with a fixed q-norm. The maximum number of interactions is again set to 2 with the LARS solver. Ordinary Kriging is applied using the Mat&eacute;rn-5/2 correlation family, ellipsoidal, and anisotropic correlation function. We used a hybrid genetic algorithm to optimize the hyperparameters.</p> <p>We benchmarked both linear and ordinary <strong>Kriging</strong>, including Mat&eacute;rn-5/2 and Gaussian correlation families and separable and ellipsoidal correlation, resulting in eight different Kriging competitors. The hyperparameters were calculated using a hybrid covariance matrix adaptation-evolution strategy optimization.</p> <p>For further details on the settings, please refer to the competitors.m file and UQLab user manuals:</p> <ul> <li>S. Marelli, N. Luethen, B. Sudret, <a href="https://www.uqlab.com/pce-user-manual">UQLab User Manual &ndash; Polynomial Chaos Expansions</a>, Report UQLab-V2.1-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>C. Lataniotis, D. Wicaksono, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/kriging-user-manual">UQLab User Manual &ndash; Kriging (Gaussian Process Modeling)</a>, Report UQLab-V2.1-105, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2024.</li> <li>R. Schoebi, S. Marelli, B. Sudret, <a href="https://www.uqlab.com/pck-user-manual">UQLab User Manual &ndash; Polynomial Chaos Kriging</a>, Report UQLab-V2.0-109, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, Switzerland, 2022.</li> </ul> <h2>Description of the results file</h2> <p>The results file contains one variable: <em>Metrics</em>. It is a Matlab structure with fields corresponding to each competitor (currently 12). Each competitor field contains data of type non-scalar struct array. The performance metrics included are RelMSE, RelRMSE, RelMAE, MAPE, Q2, and RelCVErr. Each field of Metrics.(CompetitorName) for the i-th element of the struct array contains metrics corresponding to the replicated dataset and the competitor, structured as follows:</p> <ul> <li>Metrics.(CompetitorName)(i).(MetricName)(l)<br> <ul> <li>i: dataset group,</li> <li>l: replication index.</li> </ul> </li> </ul> <p>The description of the performance measures (metrics) can be found here: <a href="https://uqworld.org/t/metamodel-performance-measures/" target="_blank" rel="noopener">https://uqworld.org/t/metamodel-performance-measures/</a>.</p> <h2>Additional files</h2> <p>We provide files in three languages (MATLAB, Python, and Julia) to showcase how to work with datasets, results, and their visualization. The files are called <em>working_with_datafiles.*</em>&nbsp;(the extension depends on the selected language).</p> <h2>Acknowledgment</h2> <p>This project was supported by the Open Research Data Program of the ETH Board under Grant number EPFL SCR0902285. The calculations were run on the Euler cluster of ETH Z&uuml;rich using the MATLAB-based UQLab software developed at the Chair of Risk, Safety and Uncertainty Quantification of ETH Z&uuml;rich.</p>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record