Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
855
datasets available to search
ShareScore release 0.9.0
Dataset results
855 results for “model system”
F-DATA: A Fugaku Workload Dataset for Job-centric Predictive Modelling in HPC Systems
<p>F-DATA is a novel workload dataset containing the data of around 24 million jobs executed on <a href="https://www.r-ccs.riken.jp/en/fugaku/">Supercomputer Fugaku</a>, over the three years of public system usage (March 2021-April 2024). Each job data contains an extensive set of features, such as exit code, duration, power consumption and performance metrics (e.g. #flops, memory bandwidth, operational intensity and memory/compute bound label), which allows for a multitude of job characteristics prediction. The full list of features can be found in the file <code>feature_list.csv</code>.</p> <p>The sensitive data appears both in anonymized and encoded versions. The encoding is based on a Natural Language Processing model and retains sensitive but useful job information for prediction purposes, without violating data privacy. The scripts used to generate the dataset are available in the<a href="https://github.com/francescoantici/F-DATA"> F-DATA GitHub repository</a>, along with a series of plots and instruction on how to load the data.</p> <p>F-DATA is composed of 38 files, with each <code>YY_MM.parquet</code> file containing the data of the jobs submitted in the month MM of the year YY. </p> <div> <div>The files of F-DATA are saved as <code>.parquet</code> files. It is possible to load such files as dataframes by leveraging the <code>pandas</code> APIs, after installing <code>pyarrow</code> (<code>pip install pyarrow</code>). A single file can be read with the following <code>Python</code> instrcutions:</div> <br> <blockquote> <div><code># Importing pandas library</code></div> <div><code>import pandas as pd</code></div> <div> </div> <div><code># Read the 21_01.parquet file in a dataframe format</code></div> <div><code>df = pd.read_parquet("21_01.parquet")</code></div> <div><code>df.head()</code></div> </blockquote> <div> </div> <div>Please cite this work as:<br><br> <div> <div>@article{antici2025fdata,</div> <div>title={F-DATA: A Fugaku Workload Dataset for Job-centric Predictive Modelling in HPC Systems},</div> <div>author={Antici, Francesco and Bartolini, Andrea and Domke, Jens and Kiziltan, Zeynep and Yamamoto, Keiji},</div> <div>journal = {Scientific Data},</div> <div>volume={12},</div> <div>pages={1321},</div> <div>year={2025},</div> <div>publisher={Nature Publishing Group},</div> <div>doi={https://doi.org/10.1038/s41597-025-05633-1}</div> <div>}</div> </div> </div> </div>
Growth Dynamics and System Models for the Restaurant Industry: Data and Analysis from Taiwanese Chains
<p><span>This dataset includes raw and curated data, system dynamics models, and feedback loop diagrams used in the study of growth dynamics in the restaurant industry, focusing on Taiwanese chains. The data supports the findings presented in the paper "The Growth Dynamics of the Restaurant Industry from Single Store to Chain Store in Taiwan: A Systems Thinking Perspective.”</span></p>
Prediction of the ground state for indenofluorene-type systems with Clar's π-sextet model
<p>This dataset contains the computational data associated with "Prediction of the ground state for indenofluorene-type systems with Clar's π-sextet model". </p>
Model outputs used in "Drivers of cross-shelf exchange in the East Auckland Current system"
<p>These simulations were developed during the work published in in Santana et al. (2023) (https://gmd.copernicus.org/articles/16/3675/2023/). The DA run is refeered as ASFUVTS (7-day assimilation window) and the NoDA run is the non-assimilative run in Santana et al., (2023). The DA run assimilates observations presented and analysed in Santana et al. (2020) (https://www.nature.com/articles/s41598-021-89222-3).</p> <p>We used the Regional Ocean Modeling System (ROMS), a primitive-equation, hydrostatic, and free-surface ocean model that solves the Reynolds-averaged form of the Navier–Stokes equations. ROMS is a fully nonlinear, finite-difference model that uses terrain-following (sigma) vertical coordinates and horizontal orthogonal or curvilinear Arakawa C grid (Shchepetkin and McWilliams, 2003, 2005; Haidvogel et al., 2008). The model domain (290 x 150) is rotated 52.14° clockwise to better resolve the NZNES and spans 332 km offshore at the widest point (near North Cape). The domain has horizontal resolution of approximately 2 km, which roughly captures coastline variability and still resolves the continental shelf and slope without large computational cost. The model has 30 vertical sigma layers and model bathymetry was interpolated from the 250 m resolution bathymetric data set built by the National Institute of Water and Atmospheric Research (NIWA - https://niwa.co.nz/our-science/oceans/bathymetry). We use a vertical discretisation scheme that increases the resolution near the surface and bottom by applying stretching function type 4 and transformation equation option 2 (Shchepetkin and McWilliams, 2005, 2009). The vertical resolution is higher at the upper 200 m (from 4 to 30 layers). On the slope (depth<1000 m), the vertical resolution is higher than 66 m and in the open ocean the thickest level is 233 m (3838.6 m depth). Baroclinic modes are resolved using a time step of 180 s, while the barotropic time step is 6 s.</p> <p>Model surface forcing is from the Japanese atmospheric 55-year reanalysis for driving ocean models (JRA55-do, Tsujino et al., 2018). A previous study demonstrated that JRA55-do had the highest correlation with observed winds in comparison to other atmospheric forcing datasets in the Southwest Pacific (Taboada et al., 2019). Atmospheric forcing fields of wind speed, net shortwave radiation, downward longwave radiation, relative humidity, temperature, rain, and pressure are specified every 3 h and used to compute the surface fluxes of stress, heat and freshwater using the bulk flux parameterisation of Fairall et al., (2003). The model uses initial and boundary conditions of SSH, temperature, salinity, and velocities from HYCOM-NCODA (Chassignet et al., 2009) versions 91.1 and 91.2 which cover period of simulations generated here. de Souza et al., (2021) analysed the performance of four global reanalysis that assimilate SSH, SST and Argo data on the New Zealand waters. The authors found that HYCOM-NCODA (8 km resolution) had higher SSH and SST variability than low-resolution (25 km) satellite surface observations and other reanalyses with similar grid-spacing. HYCOM-NCODA had velocity standard deviation similar to that observed at M4 and M5, even though GLORYS (8 km global ocean reanalysis, Lellouche et al., 2018) better represented temperature and salinity profiles in the region. Annual average discharge from several rivers are included as lateral forcing in the model. The boundary forcing is applied daily using Chapman (1985) condition for free surface, Shchepetkin condition (Mason et al., 2010) for barotropic velocities and mixed radiation-nudging (Marchesiello et al., 2001) for baroclinic velocities, temperature and salinity. A 5-day nudging coefficient is applied towards the lateral boundaries. The model aims to simulate continental shelf, slope and rise regions, including the offshore extent of the EAuC and its eddy variability. This and the upcoming studies using the simulations analysed here focus on intra-annual variability and tides are not included as forcing.</p>
Data for "SPH modelling of AGB wind morphology in hierarchical triple systems & comparison to observation of R Aql"
<div> <p>Additional material to Malfait et al. 2024, subm. "SPH modelling of AGB wind morphology in hierarchical triple systems & comparison to observation of R Aql"</p> <p>This contains input files and final output dumps of the Phantom simulations of this paper.</p> <p>The code used to perform the simulations is available at: <a href="https://github.com/danieljprice/phantom">https://github.com/danieljprice/phantom.</a></p> <p>Splash (<a href="https://github.com/danieljprice/splash">https://github.com/danieljprice/splash</a> ) and Plons (<a href="https://github.com/Ensor-code/plons">https://github.com/Ensor-code/plons</a> ) were used to create figures and plots from this data.</p> <p> </p> </div>
Modeling fluid flow in ship systems for controller tuning using an artificial neural network
<p>Dataset used to develop ANN NARX models</p>
OpenFoam model output for "A Fuel Cell Power Supply System Equipped with Artificial Gill Membranes for Underwater Applications"
<p>Dataset of numerical experiments carried out with OpenFOAM v 10 as used in the manuscript "A Fuel Cell Power Supply System Equipped with Artificial Gill Membranes for Underwater Applications" by Lucas Merckelbach and Prokopios Georgopanos.</p> <p> </p>
NUIST-Earth System Model and inputdata
<p>Fortran code and required input data of Nanjing University of Information Science and Technology (NUIST) earth system model (ocean biogeochemical version)</p>
Simulation results of two agent-based models of logistics systems
<p>The data set contains the simulation results of the two different agent-based models of logistics systems: <br> 1. A medical treatment facility (MTF) model, consisting of agents representing wounded soldiers and fixed sites representing medical facilities.<br> 2. A ship fueling (SF) simulation, consisting of agents representing fuel transport ships and fixed sites representing fuel-using bases.</p> <p>Folders in folder "MTF" are related to the MTF model.</p> <p>Files in folder "casualty_rateX" are related to the case with casualty rate of X new casualties per time step, where X = 30, 50, 70, . . . , 330.</p> <p>Each file "RunN_casualtyX_fullness.txt" has the "fullness" (defined as the ratio of the number of patients at a site to the total patient capacity of that site) of each site for each time step, where the run number N = 1, 2, 3, . . . , 100.</p> <p>Each file "RunN_casualtyX_dow.txt" has the total number of Dead Of Wounds that occur in all sites in each time step, where the run number N = 1, 2, 3, . . . , 100.</p> <p>Folders in folder "SF" are related to the SF model.</p> <p>Files in folder "siteMaxX" are related to the case with an initial (and maximum) site fuel value of X units, where X = 25, 50, 75, . . . , 200.</p> <p>Each file "assetTowedFuelHistory_siteMaxX_N.txt" has the number of towed fuel units for each asset for each time step for run number N, where N = 1, 2, 3, . . . , 100.</p> <p>Each file "assetUseFuelHistory_siteMaxX_N.txt" has the number of onboard fuel units for each asset for each time step for run number N, where N = 1, 2, 3, . . . , 100.</p> <p>Each file "siteHistory_siteMaxX_N.txt" has the number of fuel units at each site for each time step for run number N, where N = 1, 2, 3, . . . , 100.</p>
Data and analysis for "Fldgen v1.0: An Emulator with Internal Variability and Space-Time Correlation for Earth System Models"
<p>This is an archive of the raw data and analysis source code for the paper "Fldgen v1.0: An Emulator with Internal Variability and Space-Time Correlation for Earth System Models". The archive contains:</p> <ul> <li><strong>devel.Rmd : </strong>Source code for the worksheet that contains the early development and figures for the paper.</li> <li><strong>devel.html</strong> : HTML rendering of devel.Rmd</li> <li><strong>lg-ensemble-stats.Rmd </strong>: Source code for the worksheet that contains the statistical analysis described in the paper.</li> <li><strong>lg-ensemble-stats.html</strong> : HTML rendering of lg-ensemble-stats.Rmd</li> <li><strong>cc-analysis.Rmd </strong>: Analysis of the compromise conjecture raised by some readers of the paper</li> <li><strong>cc-analysis.nb.html</strong> : HTML rendering of cc-analysis.Rmd</li> <li><strong>data.tar.bz2 </strong>: Input data for the analyses above.</li> </ul> <p>The source code in this archive is written in R and requires the R runtime environment. It also uses the fldgen package, version 1.0.0, which is available at <a href="https://github.com/JGCRI/fldgen">https://github.com/JGCRI/fldgen</a></p> <p> </p>
Dataset for "Full-wave Modelling of Terahertz Frequency Plasmons in Two-Dimensional Electron Systems"
<p>Dataset underpinning the paper "Full-wave Modelling of Terahertz Frequency Plasmons in Two-Dimensional Electron Systems" by A. Dawood et al 2019 <em>J. Phys. D: Appl. Phys.</em> https://doi.org/10.1088/1361-6463/ab0ab7</p>
Dataset underpinning the paper "Terahertz plasmon resonances in two-dimensional electron systems: Modeling approaches" by S. Siaber et al 2019 Phys. Rev. Appl.
<p>Dataset underpinning the paper "Terahertz plasmon resonances in two-dimensional electron systems: Modeling approaches" by S. Siaber et al 2019 Phys. Rev. Appl.</p>
SimBench - Electrical Power System Benchmark Models
<p>SimBench (<a href="https://www.simbench.net">www.simbench.net</a>) is a research project to create a "simulation database for uniform comparison of innovative solutions in the field of network analysis, network planning and operation", which was conducted for three and a half years from 1.11.2015 to 30.04.2019. It was part of the German Federal Government's 6th Energy Research Program "Research for an Environmentally Friendly, Reliable and Affordable Energy Supply". The project was carried out by the University of Kassel, the Fraunhofer IEE, the RWTH Aachen University and the Technical University of Dortmund in accordance with the authors mentioned above. The project, coordinated by the University of Kassel, was supported by the professional advisory from six German distribution network operators: DREWAG NETZ GmbH, Energie Netz Mitte GmbH, ENSO NETZ GmbH, Netze BW GmbH, Syna GmbH and Westnetz GmbH.</p> <p>The objective of the research project SimBench is the development of a benchmark data set to support research in grid planning and operation. SimBench Grid differs from other benchmark grids under the following key points:</p> <ul> <li>Consideration of a wide range of use cases during the development of data sets</li> <li>Provision of grid data for low voltage (LV), medium voltage (MV), high voltage (HV), extra high voltage (EHV) as well as design of data sets for a suitable interconnection of a grid among different voltage levels for cross-level simulations</li> <li>Ensuring highreproducibility and comparability by providing clearly assigned load and generation time series</li> <li>Validation of the suitability of the data sets with simulation, deliberately determined grid states including suitable dimensioning of grid assets</li> </ul> <p>In total SimBench provides 13 unique electrical power system grids (EHV: 1, HV: 2, MV: 4, LV: 6). Since SimBench is enables multi-voltage simulations, this dataset includes not only 13 folder but many more. Each excerpt of the complete dataset, composed to a folder, is distinctively named by the SimBench code.</p> <p>For further information, please visit <a href="https://www.simbench.net">www.simbench.net</a> and the documentation published there.</p>
CESM1.2 simulation data for "Quantifying the cloud particle-size feedback in an Earth system model"
<p>CESM1.2-CAM5 simulation data for "Quantifying the cloud particle-size feedback in an Earth system model"</p> <p><strong>Citation: </strong>Zhu, J., & Poulsen, C. J. (2019). Quantifying the cloud particle-size feedback in an Earth system model. <em>Geophysical Research Letters</em>, <em>46</em>, 10910–10917. <a href="https://doi.org/10.1029/2019GL083829">https://doi.org/10.1029/2019GL083829</a></p> <p>Data include:</p> <p>(1) cloud liquid particle size for liquid (AREL) and ice (AREI), grid box averaged cloud liquid (CLDLIQ) and ice (CLDICE), fractional occurrence of liquid (FREQL) and ice (FREQI), and surface temperature (TS) in the preindustrial and 2xCO2 experiments; and<br> (2) the cloud feedback (lam_CLDTOT) and cloud particle-size feedback (lam_CLDEFR3L) from our PRP-based method.</p>
The Hybrid Consensus Model Based on Blockchain Self-Executing Contract for Secure E-voting System
<p><strong>Data for Review</strong></p>
Moving-block System Requirements and 9 Formal Models
<p>The package includes a set of models for a railway moving-block system:</p> <p>(a) a PDF document named Moving-block Model and Requirements.pdf, which includes a UML model of a moving-block system together with a set of requirements for the system;</p> <p>(b) a set of 9 folders, each one associated to a formal or semi-formal development tool. Each folder contains one or more models of the moving-block system from (a), developed by means of the tool. </p> <p>The models were developed using the following tool versions. Other versions may still open and verify the models.</p> <ul> <li>Simulink (2017b)</li> <li>UMC (4.7)</li> <li>UPPAAL SMC (4.1.4) </li> <li>Atelier B (4.2.1)</li> <li>ProB (1.10.2018)</li> <li>NuSMV (2.6.0)</li> <li>SPIN (6.4.9)</li> <li>CADP (2019-a)</li> <li>FDR4 (4.2.3)</li> </ul>
Moving-block System Requirements and Formal Models
<p>The package includes a set of models for a railway moving-block system:</p> <p>(a) a PDF document named Moving-block Model and Requirements.pdf, which includes a UML model of a moving-block system together with a set of requirements for the system;</p> <p>(b) a set of 9 folders, each one associated to a formal or semi-formal development tool. Each folder contains one or more model of the moving-block system from (a), developed by means of the tool. </p> <p>The models were developed using the following tool versions. Other versions may still open and verify the models.</p> <ul> <li>Simulink (2017b)</li> <li>UMC (4.7)</li> <li>UPPAAL SMC (4.1.4) </li> <li>Atelier B (4.2.1)</li> <li>ProB (1.10.2018)</li> <li>NuSMV (2.6.0)</li> <li>SPIN (6.4.9)</li> <li>CADP (2019-a)</li> <li>FDR4 (4.2.3)</li> </ul> <p> </p>
Dataset: Risk Transfer Model for Flood Risk Evolution in a Multi-reservoir System
<p>The files in this record contain data for real-time optimal flood control decision making and risk propagation under multiple uncertainties considered for publication in Water Resources Research.</p> <p> </p> <p>The files consist of:</p> <p> </p> <p>Data:</p> <ul> <li>Figure 11;</li> <li>Figure 12;</li> <li>Figure S1;</li> <li>Figure S4</li> <li>Relative prediction error</li> <li>Reservoir information</li> <li>Streamflow</li> </ul> <p>Model code:</p> <ul> <li>Calculation of entropy</li> <li>Forecasting error simulation model</li> <li>LHS</li> <li>Analytic code of transfer model</li> </ul>
Analysis of heritage stones and model wall paintings by pulsed laser excitation of Raman, laser-induced fluorescence and laser-induced breakdown spectroscopy signals with a hybrid system
<p>Laser based analysis of artworks benefits from the development of hybrid instruments where a single laser source serves to excite fluorescence, Raman and laser induced breakdown spectroscopy (LIBS) signals. Laser induced fluorescence (LIF) and Raman spectra provide information at the molecular level, while LIBS serves for identifying the elemental composition of the substrate under consideration. Studies using several excitation wavelengths on different types of materials and substrates help to develop and establish these hybrid systems for the conservation of artworks.</p>
iCESM1.2 restart file from "The Connected Isotopic Water Cycle in the Community Earth System Model Version 1"
<p>iCESM1.2 restart file at year 1850</p> <p><strong>Citation: </strong>Brady, E. C., Stevenson, S., Bailey, D., Liu, Z., Noone, D., Nusbaumer, J., … Zhu, J. (2019). The Connected Isotopic Water Cycle in the Community Earth System Model Version 1. <em>Journal of Advances in Modeling Earth Systems</em>, <em>11</em>, 2547–2566. https://doi.org/10.1029/2019MS001663</p> <p><strong>iCESM1.2 GitHub:</strong> https://github.com/NCAR/iCESM1.2</p> <p>A CLM land surface data set is included: surfdata_1.9x2.5_simyr1850_c140303.nc.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.