Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14
datasets available to search
ShareScore release 0.9.0
Dataset results
14 results for “multi-fidelity”
Multi-fidelity Generative Deep Learning Turbulent Flows
<p>Data sets for the two numerical examples in the paper <a href="https://arxiv.org/abs/2006.04731">Multi-fidelity Generative Deep Learning Turbulent Flows</a> as well as two pre-trained models. In this work, a novel multi-fidelity deep generative model is introduced for the surrogate modeling of high-fidelity turbulent flow fields given the solution of a computationally inexpensive but inaccurate low-fidelity solver. The resulting surrogate is able to generate physically accurate turbulent realizations at a computational cost magnitudes lower than that of a high-fidelity simulation. The deep generative model developed is a conditional invertible neural network, built with normalizing flows, with recurrent LSTM connections that allow for stable training of transient systems with high predictive accuracy. Data is provided from OpenFOAM LES simulations for turbulent flow over backwards step and flow around an array of cylinders.</p> <p>Data-set Files:</p> <ul> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/backward_step_testing.tar.gz?versionId=320a523f-0015-4ba3-8c6e-66733ab5a1af">backward_step_testing.tar.gz</a> - Backward step testing data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/backward_step_training.tar.gz?versionId=ac34ac15-973d-4fb8-8881-faa17eced69f">backward_step_training.tar.gz</a> - Backward step training data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinder_array_testing.tar.gz?versionId=ebbea725-c8b6-4338-978b-dc73f943552e">cylinder_array_testing.tar.gz</a> - Cylinder array testing data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinder_array_training.tar.gz?versionId=b0409c34-fb19-45c4-bbba-6968fdcfc4d8">cylinder_array_training.tar.gz</a> - Cylinder array training data.</li> </ul> <p>Pre-trained Models:</p> <ul> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/bstepWorkspace400.zip">bstepWorkspace400.zip</a> - Backward step pre-trained model.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinderWorkspace400.zip">cylinderWorkspace400.zip</a> - Cylinder array pre-trained model.</li> </ul> <p> </p>
Reference dataset of multi-objective and multi-fidelity optimization in laser-plasma acceleration
<p>This repository contains a dataset used for the article "<em>Multi-objective and multi-fidelity Bayesian optimization of laser-plasma acceleration</em>" (<a href="https://arxiv.org/abs/2210.03484">arXiv:2210.03484</a>). The dataset consists of 2443 FBPIC particle-in-cell simulations of a laser wakefield accelerator that were selected using a Bayesian optimizer. The goal of the optimization was to perform multi-objective multi-fidelity optimization of electron beam parameters. The dataset contains simulations of different resolutions, accordingly with differing fidelities. The typical runtime at lowest (highest) resolution is approximately 1 (90) minutes.</p> <p>In the dataset we have <em>train_x </em>and <em>train_obj </em>numpy arrays with dimensions <em>(n,5)</em> and<em> (n,3)</em>, respectively. Here <em>n</em> is the number of FBPIC simulations. The five columns in <em>train_x </em>are [plasma density, upramp length, laser focus, downramp length, fidelity]. The fidelity parameter controls the resolution and hence the runtime of the simulation. The three columns in the <em>train_obj </em>are the [total charge, distance of median to target energy, bandwidth of electron beams]. For the distance, the target energy is fixed to 300 MeV and for the bandwidth is defined by the median absolute deviation around the median. The two columns have negative values since the optimizer assumes a maximization of all objectives while the distance and bandwidth in this study were being minimized.</p> <p>The different folders contain data of different kind of single and multi-objectives that were used to produce figures 2, 3, 5 in the associated paper. For more details please see the referred article. The folder "combined" contains the data of all simulations together and is most suitable for (5D x 3D) surrogate model generation.</p>
Materials Science Optimization Benchmark Dataset for Multi-Objective, Multi-Fidelity Optimization of Hard-Sphere Packing Simulations
<p>Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a “Turing test” of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In the fields of materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 494498 random hard-sphere packing simulations representing 206 CPU days worth of computational overhead were performed across nine input parameters with linear constraints and two discrete fidelities each with continuous fidelity parameters and results were logged to a free-tier shared MongoDB Atlas database. Two core tabular datasets resulted from this study: 1. a failure probability dataset containing unique input parameter sets and the estimated probabilities that the simulation will fail at each of the two steps, and 2. a regression dataset mapping input parameter sets (including repeats) to particle packing fractions and computational runtimes for each of the two steps. These two datasets are used to create a surrogate model as close as possible to running the actual simulations by incorporating simulation failure and heteroskedastic noise. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This is in contrast with a more traditional approach that imposes a-priori assumptions such as Gaussian noise e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.</p> <p>For usage instructions, see https://matsci-opt-benchmarks.readthedocs.io/.</p>
Multi-fidelity modelling of shark skin denticle flows: Insights into drag generation mechanisms
<p>We investigate the flow over smooth (non-ribletted) shark skin denticles in an open-channel flow using Direct Numerical Simulation (DNS) and two Reynolds Averaged Navier-Stokes (RANS) closures. Large peaks in pressure and viscous drag are observed at the denticle crown edges, where they are exposed to high-speed fluid which penetrates between individual denticles, increasing shear and turbulence. Strong lift forces lead to a positive spanwise torque acting on individual denticles, potentially encouraging bristling if the denticles were not fixed. However, DNS predicts that denticles ultimately increase drag by 58 % compared to a flat plate.</p> <p>Good predictions of drag distributions are obtained by RANS models, although an underestimation of turbulent kinetic energy production leads to an underprediction of drag. Nevertheless, RANS methods correctly predict trends in the drag data and the regions contributing most to viscous and pressure drag. Subsequently, RANS models are used to investigate the dependence of drag on the flow blockage ratio (boundary layer to roughness height ratio), finding that the drag increase due to denticles is halved when the blockage ratio δ /h is increased from 14 to 45. Our results provide an integrated understanding of the drag over non-ribletted denticles, enabling existing diverse drag data to be explained.</p>
Materials Science Optimization Benchmark Dataset for High-dimensional, Multi-objective, Multi-fidelity Optimization of CrabNet Hyperparameters
Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a "Turing test" of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 173219 quasi-random hyperparameter combinations were generated across 23 hyperparameters and used to train CrabNet on the Matbench experimental band gap dataset. The results were logged to a free-tier shared MongoDB Atlas dataset. This study resulted in a regression dataset mapping hyperparameter combinations (including repeats) to MAE, RMSE, computational runtime, and model size for CrabNet model trained on the Matbench experimental band gap benchmark task1. This dataset is used to create a surrogate model as close as possible to running the actual simulations by incorporating heteroskedastic noise. Failure cases for bad hyperparameter combinations were excluded via careful construction of the hyperparameter search space, and so were not considered as was done in prior work. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This contrasts with a more traditional approach that imposes a-priori assumptions such as Gaussian noise, e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.
Multi-fidelity Gaussian Process Emulation for Atmospheric Radiative Transfer Models
<p>This repository contains several datasets of spectral atmospheric transfer functions (i.e. path radiance, transmittances, spherical albedo) simulated with MODTRAN6 atmospheric radiative transfer model. The simulations are stored in hdf5 files using the Atmospheric Look-up table Generator (ALG) toolbox (<a href="https://doi.org/10.5194/gmd-13-1945-2020">https://doi.org/10.5194/gmd-13-1945-2020</a>). Each dataset has an associated .xml file that includes the configuration of ALG/MODTRAN6 executions. All datasets include the input atmospheric/geometric variables that are summarized in the following table. Each dataset file has a random distribution (based on latin hypercube sampling) these input variables with varying number of points (e.g. train500.h5 contains 500 samples). The <em>reference </em>dataset contains 10000 samples and was used as reference for evaluating Gaussian Processes emulators.</p> <table> <tbody><tr> <th>Input Variables</th> <th>Units</th> <th>Min</th> <th>Max</th> </tr> </tbody><tbody> <tr> <td>O3 column concentration</td> <td>atm-cm</td> <td>0.25</td> <td>0.45</td> </tr> <tr> <td>Columnar Water Vapor</td> <td>g/cm2</td> <td>0.2</td> <td>4</td> </tr> <tr> <td>Aerosol Optical Thickness</td> <td>-</td> <td>0.04</td> <td>0.6</td> </tr> <tr> <td>Asymmetry parameter</td> <td>-</td> <td>0.5</td> <td>0.85</td> </tr> <tr> <td>Angstrom exponent</td> <td>-</td> <td>0.1</td> <td>2</td> </tr> <tr> <td>Single Scattering Albedo</td> <td>-</td> <td>0.8</td> <td>1</td> </tr> <tr> <td>Surface elevation</td> <td>km</td> <td>0</td> <td>2.5</td> </tr> <tr> <td>Solar Zenith Angle</td> <td>deg</td> <td>0</td> <td>70</td> </tr> <tr> <td>Relative Zenith Angle</td> <td>deg</td> <td>0</td> <td>180</td> </tr> </tbody> </table> <p> </p>
Multi-fidelity modelling of shark skin denticle flows: Insights into drag generation mechanisms
Open the record for dataset details and reuse information.
Finding Efficient Trade-offs in Multi-Fidelity Response Surface Modeling: Generated data files and figures
<p>All data files and figures generated for the paper "Finding Efficient Trade-offs in Multi-Fidelity Response Surface Modeling".</p> <p>The code used to generate this is archived at <a href="https://doi.org/10.5281/zenodo.6123254">zenodo.org/record/6123254</a></p>
Codes and dataset used in the manuscript entitled "Quantifying time-variant travel time distribution by multi-fidelity model in hillslope under nonstationary hydrologic conditions"
<p>This contains the codes and dataset for the manuscript entitled "Quantifying time-variant travel time distribution by multi-fidelity model in hillslope under nonstationary hydrologic conditions". Detailed information about the dataset is described in the Readme.txt file.</p>
Multi-fidelity reduced-order surrogate modeling
<p>Training and testing datasets used for the experiments in <a href="https://arxiv.org/abs/2309.00325">Multi-fidelity reduced-order surrogate modeling</a></p>
Dataset for paper: Multi-Fidelity Surrogate Modelling of Wall Mounted Cubes
<p>Dataset including samples of LES and RANS data of the flow around tandem wall mounted cubes.</p> <p>Samples are taken as lines along the centre of the domain and a number of slices.</p> <p>This data is used in the paper "Multi-Fidelity Surrogate Modelling of Wall Mounted Cubes".</p> <p>The data is used by code found in the repository https://github.com/admole/Multi-Fidelity-Surrogate/tree/FTaC</p>
Materials Science Optimization Benchmark Dataset for Multi-fidelity Hard-sphere Packing Simulations
Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a "Turing test" of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In the fields of materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 438371 random hard-sphere packing simulations representing 279 CPU days worth of computational overhead were performed across nine input parameters with linear constraints and two discrete fidelities each with continuous fidelity parameters and results were logged to a free-tier shared MongoDB Atlas database. Two core tabular datasets resulted from this study: 1. a failure probability dataset containing unique input parameter sets and the estimated probabilities that the simulation will fail at each of the two steps, and 2. a regression dataset mapping input parameter sets (including repeats) to particle packing fractions and computational runtimes for each of the two steps. These two datasets can be used to create a surrogate model as close as possible to running the actual simulations by incorporating simulation failure and heteroskedastic noise. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This is in contrast with a more traditional approach that imposes a-priori assumptions such as Gaussian noise e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.
On the characteristics of the wake of a wind turbine undergoing large motions caused by a floating structure: an insight based on experiments and multi-fidelity simulations from the OC6 Phase III Project - Supplementary material
<p>This link stores the supplemenary material for the paper: "On the characteristics of the wake of a wind turbine undergoing large motions caused by a floating structure: an insight based on experiments and multi-fidelity simulations from the OC6 Phase III Project" published on Wind Energy Science. The pdf file contains the plots for all the investigated metrics analysed in this work, which could not be reported in the manuscript due to space constraints. Further details about these results can be found in the original publication. </p>
ARC Code TI: Multi-Fidelity Simulator (MFSim)
Multi-Fidelity Simulator, MFSim is a pluggable framework for creating an air traffic flow simulator at multiple levels of fidelity. The framework is designed to allow low-fidelity simulations of the entire US Airspace to be completed very quickly (on the order of seconds). The framework allows higher-fidelity plugins to be added to allow higher-fidelity simulations to occur in certain regions of the airspace concurrently with the low-fidelity simulation of the full airspace.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.