Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
70
datasets available to search
ShareScore release 0.7.1
Dataset results
70 results for “HPC”
Performance Measurement Dataset of the HPC Benchmarks FASTEST, Kripke, and RELeARN for Cost-Effective Modeling Analysis with Extra-P
<p>Performance Measurement Dataset of the HPC Benchmarks FASTEST, Kripke, RELeARN for Scalability Studies with Extra-P. This data was used to analyze cost-effective modeling approaches presented in the IPDPS 2020 paper "Learning Cost-Effective Sampling Strategies for Empirical Performance Modeling".</p>
Processed files used for OQ based tsunami loss modelling using HPC based inundation and emulators(Part-IV OQ Simulation and Emulation Hazard Dataset)
<div> <div>This dataset is related to the main Zenodo repository: https://doi.org/10.5281/zenodo.13738078</div> <br> <div>This dataset contains some of the processed OpenQuake files covering tsunami inundation depth hazard using simulation results and emulation results discussed in the article - "Towards Using Machine Learning Emulation for Probabilistic Inundation Mapping: Multiple Earthquake Sources and Near-field Effects with project repo - https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators.</div> <div> </div> <div>The risk calculation and procedure is available in the main repo: <a href="https://github.com/naveenragur/ML4SicilyTsunami/tree/main-dev/risk">https://github.com/naveenragur/ML4SicilyTsunami/tree/main-dev/risk </a></div> <div> </div> <br> <div>The processed files for the test locations of Catania(CT) are provided in compressed gzip files(.gz):</div> <br> <div>Processed hazard information used to prepare and run OQ event based risk analysis are as below,</div> <div> -<strong>hazard.tar.gz </strong>- preliminary numpy files with sitcol, eventid and hazard magnitude info</div> <div> -<strong>loss.tar.gz</strong> -final hdf5 files with both event hazard and site info</div> <br> <div>The filename follows the nomenclature with:</div> </div> <div>ML4SicilyTsunami/risk/loss/tsunami_prob_892.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_1658.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_3454.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_7071.hdf5<br>ML4SicilyTsunami/risk/loss/tsunami_prob_true.hdf5</div> <div><br> <div> <div>More information on the attached readme, see project structure and code is available at:</div> <div><strong>https://github.com/naveenragur/ML4SicilyTsunami/tree/ptha_emulators</strong></div> <div><strong>https://github.com/naveenragur/ML4SicilyTsunami/tree/main-dev/risk</strong></div> <div><strong>https://github.com/naveenragur/OQ-Tsunami</strong></div> </div> </div>
Power, performance and system measures of HPC benchmarks on multiple hardware
Open the record for dataset details and reuse information.
HPC geophysical electromagnetics: a synthetic VTI model with complex bathymetry
<p>Castillo-Reyes, O., de la Puente, J., Cela, E. J.M. (2022) HPC geophysical electromagnetics: a synthetic VTI model with complex bathymetry. Submitted to Energies Journal</p>
Companion data of Multi-Phase Task-Based HPC Applications: Quickly Learning how to Run Fast
<p>This is the companion data repository for the paper entitled <strong>Multi-Phase Task-Based HPC Applications: Quickly Learning how to Run Fast</strong> by Lucas Leandro Nesi, Lucas Mello Schnorr, and Arnaud Legrand. The manuscript has been accepted for publication in the <a href="https://www.ipdps.org/ipdps2022/2022-organization.html">IPDPS 2022</a>.</p>
Artifact of the paper: Light-weight prediction for improving energy consumption in HPC platforms
<p>Please refer to the <a href="../records/11208389/files/artifact-overview.pdf?download=1&preview=1" target="_blank" rel="noopener">artifact-overview.pdf</a> file in this dataset for instructions to reproduce the experiments we have conducted for this article, or for more context about the article.</p>
Companion data of a Systematic Mapping Study of Programming Languages for Data-Intensive HPC Applications
<p>As the current existing literature on the topic of HPC is very dispersed, we performed a Systematic Mapping Study (SMS) in the context of the European COST Action cHiPSet. This literature study maps characteristics of various programming languages for data-intensive HPC applications, including category, typical user profiles, effectiveness, and type of articles.</p> <p>We organised the SMS in two phases. In the first phase, relevant articles are identified employing an automated keyword-based search in eight digital libraries. This lead to an initial sample of 420 papers, which was then narrowed down in a second phase by human inspection of article abstracts, titles and keywords to 152 relevant articles published in the period 2006--2018. The analysis of these articles enabled us to identify 26 programming languages referred to in 33 of relevant articles. This document is the data companion for a paper published elsewhere and presents a detailed list of the selected papers. Besides, the document also presents the form of our questionnaire-based survey. </p> <p>We also include the filled in questionnaires and raw data of the referred survey. To validate the SMS results we conducted a survey (in November 2018) with 28 HPC experts involved in the cHiPSet COST action to which we added, in October 2019, 29 HPC experts which were not involved in that COST action. Participants were recruited through convenience sampling, and contacted directly by the authors. In total, we received 57 filled survey forms.</p>
Companion data of a Systematic Mapping Study of Programming Languages for Data-Intensive HPC Applications
<p>As the current existing literature on the topic of HPC is very dispersed, we performed a Systematic Mapping Study (SMS) in the context of the European COST Action cHiPSet. This literature study maps characteristics of various programming languages for data-intensive HPC applications, including category, typical user profiles, effectiveness, and type of articles.</p> <p>We organised the SMS in two phases. In the first phase, relevant articles are identified employing an automated keyword-based search in eight digital libraries. This lead to an initial sample of 420 papers, which was then narrowed down in a second phase by human inspection of article abstracts, titles and keywords to 152 relevant articles published in the period 2006--2018. The analysis of these articles enabled us to identify 26 programming languages referred to in 33 of relevant articles. This document is the data companion for a paper published elsewhere and presents a detailed list of the selected papers. Besides, the document also presents the form of our questionnaire-based survey. </p> <p>We also include the filled in questionnaires and raw data of the referred survey. To validate the SMS results we conducted a survey (in November 2018) with 28 HPC experts involved in the cHiPSet COST action to which we added, in October 2019, 29 HPC experts which were not involved in that COST action. Participants were recruited through convenience sampling, and contacted directly by the authors. In total, we received 57 filled survey forms.</p>
Artifact of the paper: Scheduling with lightweight predictions in power-constrained HPC platforms
<p>Please refer to the <a href="https://zenodo.org/records/13961003/files/artifact-overview.pdf?download=1&preview=1">artifact-overview.pdf</a> file in this dataset for instructions to reproduce the experiments we have conducted for this article, or for more context about the article.</p>
Dataset for IOMax: Maximizing Out-of-Core I/O Analysis Performance on HPC Systems
<p>I/O analysis is an essential task for improving the performance of scientific applications on high-performance computing (HPC) systems. However, current analysis tools, which often use data drilling techniques (iterative exploration for deeper insights), treat every query independently and do not optimize column data for data-slicing (extracting specific data subsets), resulting in subpar querying performance. In this paper, we designed IOMax, a tool for efficient data drilling analysis on large-scale I/O traces. IOMax utilizes a novel query optimization technique to improve the query performance by 8.6x while reducing the memory footprint required for analysis by 11x. Additionally, it employs data transformation techniques to improve data-slicing performance by up to 11.4x. In conclusion, IOMax optimizes I/O analysis for scientific workflows on the Lassen supercomputer, resulting in up to 7x improvement.</p> <p>This dataset contains optimized data that was used in the associated publication.</p>
Introduction to HPC: molecular dynamics simulations with GROMACS: log files
<p>Introduction to HPC: molecular dynamics simulations with GROMACS: log files corresponding to the exercises 1.X 2.X 3.X</p>
Introduction to HPC: molecular dynamics simulations with GROMACS: output files - Devana
<p> Introduction to HPC: molecular dynamics simulations with GROMACS: log files corresponding to the exercises 1.1, 1.2 2.1 3.1 and 3.2</p>
data artifact for "Keeping It Real: Why HPC Data Services Don't Achieve I/O Microbenchmark Performance"
<p>This is the data artifact associated with the PDSW 20 workshop paper submission entitled "Keeping It Real: Why HPC Data Services Don’t Achieve I/O Microbenchmark Performance" in accordance with the PDSW 20 artifact submission guidelines.</p>
Maximizing Data Utility for HPC Python Workflow Execution
<p>Large-scale HPC workflows are increasingly implemented in dynamic languages such as Python, which allow for more rapid development than traditional techniques. However, the cost of executing Python applications at scale is often dominated by the distribution of common datasets and complex software dependencies. As the application scales up, data distribution becomes a limiting factor that prevents scaling beyond a few hundred nodes. To address this problem, we present the integration of Parsl (a Python-native parallel programming library) with TaskVine (a data-intensive workflow execution engine). Instead of relying on a shared filesystem to provide data to tasks on demand, Parsl is able to express advance data needs to TaskVine, which then performs efficient data distribution at runtime. This combination provides a performance speedup of 1.48x over the typical method of on-demand paging from the shared filesystem, while also providing an average task speedup of 1.79x with 2048 tasks and 256 nodes.</p>
PM100: A Job Power Consumption Dataset of a Large-Scale HPC System
<p>The dataset is a collection of jobs extracted from the job_table data of M100 (<a href="https://doi.org/10.5281/zenodo.7588815">https://doi.org/10.5281/zenodo.7588815</a>), a collection of data extracted from a Tier-0 supercomputer hosted at CINECA (Marconi100, <a href="https://www.hpc.cineca.it/hardware/marconi100">https://www.hpc.cineca.it/hardware/marconi100</a>). The original job data present in M100 are filtered out by considering only the jobs running exclusively on the resources. Each job entry included in PM100 contains the power consumption of the job recorded at Node level, CPU level and Memory level. The final dataset contains 231116 jobs, executed on Marconi100 between May and October 2020. </p><p>The dataset is stored as a parquet file, where each entry contains the information on a job execution. </p><p>The structure of the data, as well as the code to generate them, is contained in the official GitHub repository of the project: <a href="https://github.com/francescoantici/PM100-data/">https://github.com/francescoantici/PM100-data/</a>.</p>
Performance Measurement Datasets of the HPC Benchmarks LAMMPS, MiniFE, LULESH for Hardware Counter Variance Analysis
Open the record for dataset details and reuse information.
Data publication: Simulation of macroscopic boundary value problems using phenomenological material model for steel fiber reinforced high performance concrete (HPC)
<p>This data set contains all necessary inputs for the Simulation of macroscopic boundary value problems using phenomenological material model, including boundary conditions, material parameters and numerical results. The discretization is realized in terms of the finite element method. </p>
F-DATA: A Fugaku Workload Dataset for Job-centric Predictive Modelling in HPC Systems
<p>F-DATA is a novel workload dataset containing the data of around 24 million jobs executed on <a href="https://www.r-ccs.riken.jp/en/fugaku/">Supercomputer Fugaku</a>, over the three years of public system usage (March 2021-April 2024). Each job data contains an extensive set of features, such as exit code, duration, power consumption and performance metrics (e.g. #flops, memory bandwidth, operational intensity and memory/compute bound label), which allows for a multitude of job characteristics prediction. The full list of features can be found in the file <code>feature_list.csv</code>.</p> <p>The sensitive data appears both in anonymized and encoded versions. The encoding is based on a Natural Language Processing model and retains sensitive but useful job information for prediction purposes, without violating data privacy. The scripts used to generate the dataset are available in the<a href="https://github.com/francescoantici/F-DATA"> F-DATA GitHub repository</a>, along with a series of plots and instruction on how to load the data.</p> <p>F-DATA is composed of 38 files, with each <code>YY_MM.parquet</code> file containing the data of the jobs submitted in the month MM of the year YY. </p> <div> <div>The files of F-DATA are saved as <code>.parquet</code> files. It is possible to load such files as dataframes by leveraging the <code>pandas</code> APIs, after installing <code>pyarrow</code> (<code>pip install pyarrow</code>). A single file can be read with the following <code>Python</code> instrcutions:</div> <br> <blockquote> <div><code># Importing pandas library</code></div> <div><code>import pandas as pd</code></div> <div> </div> <div><code># Read the 21_01.parquet file in a dataframe format</code></div> <div><code>df = pd.read_parquet("21_01.parquet")</code></div> <div><code>df.head()</code></div> </blockquote> <div> </div> <div>Please cite this work as:<br><br> <div> <div>@article{antici2025fdata,</div> <div>title={F-DATA: A Fugaku Workload Dataset for Job-centric Predictive Modelling in HPC Systems},</div> <div>author={Antici, Francesco and Bartolini, Andrea and Domke, Jens and Kiziltan, Zeynep and Yamamoto, Keiji},</div> <div>journal = {Scientific Data},</div> <div>volume={12},</div> <div>pages={1321},</div> <div>year={2025},</div> <div>publisher={Nature Publishing Group},</div> <div>doi={https://doi.org/10.1038/s41597-025-05633-1}</div> <div>}</div> </div> </div> </div>
Introduction to HPC: molecular dynamics simulations with GROMACS: input files
<p>Introduction to HPC: molecular dynamics simulations with GROMACS: input files</p>
Efficacy and Safety of HPC-03 for Postmenopausal Symptom
ClinicalTrials.gov study NCT03061799. IPD Sharing: NO. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.