Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
70
datasets available to search
ShareScore release 0.7.1
Dataset results
70 results for “HPC”
Supplementary Materials for "On the Scalability of Data Reduction Techniques in Current and Upcoming HPC Systems from an Application Perspective"
<p>Supplementary materials with all used benchmark scripts, plot scripts, benchmark results and PIConGPU example data for the submission to "The 1st International Workshop on Data Reduction for Big Scientific Data (DRBSD-1)" held in conjunction with ISC 2017 in Frankfurt, Germany.</p>
Data set for anomaly detection on a HPC system
<p>This data set contains the data collected on the DAVIDE HPC system (CINECA & E4 & University of Bologna, Bologna, Italy) in the period March-May 2018.</p> <p>The data set has been used to train a autoencoder-based model to automatically detect anomalies in a semi-supervised fashion, on a real HPC system.</p> <p>This work is described in:</p> <p>1) "Anomaly Detection using Autoencoders in High Performance Computing Systems", <a href="https://arxiv.org/search/cs?searchtype=author&query=Borghesi%2C+A">Andrea Borghesi</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Bartolini%2C+A">Andrea Bartolini</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Lombardi%2C+M">Michele Lombardi</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Milano%2C+M">Michela Milano</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Benini%2C+L">Luca Benini,</a> IAAI19 (proceedings in process) -- https://arxiv.org/abs/1902.08447</p> <p>2) "Online Anomaly Detection in HPC Systems", <a href="https://arxiv.org/search/cs?searchtype=author&query=Borghesi%2C+A">Andrea Borghesi</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Libri%2C+A">Antonio Libri</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Benini%2C+L">Luca Benini</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Bartolini%2C+A">Andrea Bartolini, </a>AICAS19 (proceedings in process) -- https://arxiv.org/abs/1811.05269</p> <p>See the git repository for usage examples & details --> https://github.com/AndreaBorghesi/anomaly_detection_HPC</p>
A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p><span><span><span><span>We demonstrate a resilient workflow enabled by the LEXIS Platform, running a time- and safety-critical biomedical simulation of virtual stent placement in intracranial arteries using the HemoFlow application. The workflow, as captured on the video, gracefully handles failures of single computing steps or entire computing systems and thus lends itself to urgent computing applications. <br><br><span><span>The concept of this workflow has potential for realising ab-initio computational biomedical simulations which can provide live, targeted guidance to surgeons.</span></span></span></span></span></span></p>
Additional Artifacts - Supplements to: A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p>In this dataset, we have collected supplementary artifacts to support an understanding of the workflow presented in the submission cited (see related identifiers).</p> <p>These artifacts are (cf. README.md in the main folder of the tar.gz archive):</p> <p>A1: modified HemoFlow code (cf. https://github.com/gzavo/hemoflow) for our workflow experiments (subfolder "hemoflowcfd");<br>A2: workflow descriptions in python for Apache Airflow (subfolder "workflow");<br>A3: inputs (.xml/.npz) and output (.txt) for the example (subfolder "case").</p> <p> </p>
Dataset and scripts for the paper with title Evaluating Programming Models for the HPC GPU Ecosystem
<p>Dataset and scripts for the paper with title Evaluating Programming Models for the HPC GPU Ecosystem</p>
HPC-JEEP: Energy Usage on ARCHER2 and the DiRAC COSMA HPC services dataset
<p>This package contains the data and tools used to analyse the energy use on the ARCHER2 and DiRAC COSMA UK HPC facilities. This analysis was performed as part of the HPC-JEEP project. HPC-JEEP is funded by the UKRI DRI Net Zero Scoping project.</p>
Antarex HPC Fault Dataset
<p>The Antarex dataset contains trace data collected from the homonymous experimental HPC system located at ETH Zurich while it was subjected to fault injection, for the purpose of conducting machine learning-based fault detection studies for HPC systems. Acquiring our own dataset was made necessary by the fact that commercial HPC system operators are very reluctant to share trace data containing information about faults in their systems.</p> <p>In order to acquire data, we executed benchmark applications and at the same time injected faults in the system at specific times via dedicated programs, so as to trigger anomalies in the behaviour of the applications. A wide range of faults is covered in our dataset, from hardware faults, to misconfiguration faults, and finally to performance anomalies cause by interference from other processes. This was achieved through the FINJ fault injection tool, developed by the authors.</p> <p>The dataset contains two types of data: one type of data refers to a series of CSV files, each containing a set of system performance metrics sampled through the LDMS HPC monitoring framework. Another type refers to the log files detailing the status of the system (i.e., currently running benchmark applications or injected fault programs) at each time point in the dataset. Such a structure enables researchers to perform a wide range of studies on the dataset. Moreover, since we collected the dataset by streaming continuous data, any study based on it will easily be reproducible on a real HPC system, in an online way. The dataset is divided in two parts: the first includes only the CPU and memory-related benchmark applications and fault programs, while the second is strictly hard drive-related. We executed each part in both single-core and multi-core variants, resulting in a total of 4 dataset blocks for 32 days of data acquisition, and 20GB of uncompressed data.</p> <p>For a detailed analysis on the structure and features of the Antarex dataset, please refer to the research paper "Online Fault Classification in HPC System through Machine Learning", by Netti et al. Additional details can be found in the research paper "FINJ: a Fault Injection Tool for HPC System" by Netti et al., whereas all source code can be found on the GitHub repository of the FINJ tool.</p> <p>When using this dataset, please cite the two reference papers above as follows:</p> <p>" Netti A., Kiziltan Z., Babaoglu O., Sîrbu A., Bartolini A., Borghesi A. (2019) FINJ: A Fault Injection Tool for HPC Systems. In: Mencagli G. et al. (eds) Euro-Par 2018: Parallel Processing Workshops. Euro-Par 2018. Lecture Notes in Computer Science, vol 11339. Springer, Cham"</p> <p>" Netti A., Kiziltan Z., Babaoglu O., Sîrbu A., Bartolini A., Borghesi A. (2019) Online Fault Classification in HPC Systems through Machine Learning. arXiv:1810.11208"</p>
HPC-JEEP: Energy-based charging on the ARCHER2 HPC service dataset
<p>This package contains the data and tools used to calculate and analyse an approach to energy-based charging on the ARCHER2 UK HPC facility. This analysis was performed as part of the <a href="https://zenodo.org/record/6787599/">HPC-JEEP project</a>. HPC-JEEP is funded by the <a href="https://net-zero-dri.ceda.ac.uk/">UKRI DRI Net Zero Scoping project</a>.</p>
HPC Production Trace
<p>Job trace includes 14M jobs from a production high performance computing cluster consisting of 14,376 cores. Each job entry includes its submission time, user ID, maximum running time limit, requested number of cores and memory, and running time. We have formatted the trace in hierarchical data format (hdf) format).</p> <p>## Important Data Fields</p> <p>* *userid*: The user identification number. (Integer)</p> <p>* *wallclock_runtime_sec*: Actual job runtime in seconds. (Integer)</p> <p>* *wallclock_limit_sec*: Maximum runtime limit specified by the user. (Integer)</p> <p>* *num_cores*: Number of CPU requested for the job. (Integer)</p> <p>* *total_MB_req*: Memory in MB requested for the job. (Integer)</p> <p>* *status*: Status of the job (String)</p> <p>* *time*: Timestamp when job queued. (Timestamp)</p>
Companion for "Understanding Distributed Deep Learning Performance by Correlating HPC and Machine Learning Measurements"
<p>This is the Companion Material for the paper “Understanding Distributed Deep Learning Performance by Correlating HPC and Machine Learning Measurements”, by Ana Luisa Veroneze Solórzano and Lucas Mello Schnorr. The manuscript was approved for publication in the <a href="https://www.isc-hpc.com/research-papers-2022.html">ISC High Performance 2022</a> for the Research Papers session. A public companion is also availabl in GitLab: <a href="https://gitlab.com/anaveroneze/isc2022-companion/">https://gitlab.com/anaveroneze/isc2022-companion</a>.</p> <p> </p>
Optimising bioelectrochemical systems using Hartree HPC Datasets
<p>A dataset containing simulation results for the optimising bioelectrochemical systems project.</p> <p>This contains steady-state data packed within TAR files. This data has been extracted and summarised within the CSV files.</p> <p>Data was gathered from adapting the mathematical model for simulation of bioelectrochemical systems published in Day et al 2022, this code will be released upon the submission of JDays thesis in late 2022. Steady-state data is saved as compressed NumPy files and can be extracted using python.</p> <p> </p>
Data publication: Virtual experiments for steel fiber reinforced high performance concrete (HPC)
<p>This data set contains all necessary inputs for the virtual experiments using an ellipsodal RVE for steel fiber reinforced high performance concrete (HPC), including discretization data, boundary conditions, material parameters and numerical results. The discretization is realized in terms of the finite element method. </p>
Comparability and Reproducibility in HPC Applications' Energy Consumption Characterization
<p>The computational power of HPC systems continues to grow, and improving their energy efficiency is a critical issue for the field in the face of climate change and energy crises. One major aspect of energy optimization lies in the applications run on the systems themselves. In this work, we are looking into comparing energy consumption between different systems using a characterization process based on the recent energy characterization paper as a reference and starting point for other data centers to assess their application’s energy patterns. We demonstrated that we could use the methods from the starting paper, replicate the findings, and extend the work to more applications and more systems. Our work acts as a proof of concept for a repository of HPC applications’ energy patterns in our future work.<br><br>This is the collection of jobscripts, data, and python scripts used in the paper.</p>
Generic and ML Workloads in an HPC Datacenter
<p>Updated Version of the <a title="previous upload" href="../records/13625495">previous upload</a>, adjusts node timestamps lacking behind at the beginning of the data collection.</p> <p>This archive contains hardware and workload traces from SURF Lisa, a Dutch datacenter consisting of 338 nodes, used by universities and researchers for various jobs. Around 85% of the nodes are equipped only with CPUs, handling generic compute-heavy workloads, the other 15% come with additional GPUs, serving as accelerators for Machine Learning (ML) jobs. Individual node hardware configurations are listed in `node_hardware_info.parquet`.</p> <p>Jobs within Lisa are submitted over the SLURM scheduler, where we logged job start and end time, resource allocation, and exit state for roughly 10 months (December 2021 to November 2022). This data saved in `slurm_table_cleaned.parquet`.</p> <p>Addidionally, we provide detailed Prometheus monitoring logs from all nodes over a timespan of 5 months (June 2022 to November 2022) in `prom_table_cleaned.parquet`. These logs contain over 90 attributes, including CPU/GPU power and temperatures, network I/O, memory and storage usage, and many more. These metrics are sampled at 30s intervals, resulting in a total of almost 130 million records across all nodes.</p> <p>Finally, job and node data are provided as a joined dataset in `prom_slurm_joined.parquet` for their 4 months of overlapping timespan. This combined data can provide more insights into the resource consumption and performance patterns of jobs.</p> <p>We conducted detailed analysis of this data where we specifically looked at the different characteristics of generic vs. ML workloads in a heterogeneous HPC environment. The pre-print of our analysis work can be found on <a href="https://arxiv.org/abs/2409.08949">arXiv</a>. Our code used for evaluation can be found on <a href="https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization">GitHub</a>.<br><br></p> <table> <tbody> <tr> <th>Dataset Name</th> <th>Explanation</th> </tr> <tr> <td>slurm_table_cleaned.parquet</td> <td>Job data collected by SLURM</td> </tr> <tr> <td>prom_table_cleaned.parquet</td> <td>Node data collected by Prometheus</td> </tr> <tr> <td>prom_slurm_joined.parquet</td> <td>Joined Job and Node dataset</td> </tr> <tr> <td>node_hardware_info.parquet</td> <td>Hardware configurations of each node</td> </tr> </tbody> </table>
Dataset Artifact for Prodigy: Towards Unsupervised Anomaly Detection in Production HPC Systems
<p>The dataset contains a small set of application runs from Eclipse supercomputer. The applications run with and without synthetic HPC performance anomalies. More detailed information regarding synthetic anomalies can be found at: https://github.com/peaclab/HPAS.</p> <p>We have chosen four applications, namely LAMMPS, sw4, sw4Lite, and ExaMiniMD, to encompass both real and proxy applications. We have executed each application five times on four compute nodes without introducing any anomalies. To showcase our experiment, we have specifically selected the "memleak" anomaly as it is one of the most commonly occurring types. Additionally, we have also executed each application five times with the chosen anomaly. The dataset we have collected consists of a total of 160 samples, with 80 samples labeled as anomalous and 80 samples labeled as healthy. For the details of applications please refer to the paper.</p> <p>The applications were run on Eclipse, which is situated at Sandia National Laboratories. Eclipse comprises 1488 compute nodes, each equipped with 128GB of memory and two sockets. Each socket contains 18 E5-2695 v4 CPU cores with 2-way hyperthreading, providing substantial computational power for scientific and engineering applications.</p>
Dataset from Paper: What does Power Consumption Behavior of HPC Jobs Reveal?
<p>The dataset in the tarball was used as job- and power-trace input for the paper "What does Power Consumption Behavior of HPC Jobs Reveal?", published at the International Parallel and Distributed Processing Symposium 2020 (IPDPS'20) in New Orleans, Louisiana.</p> <p>For more details on files, clusters, and data format, please see the README file in the archive.</p>
Exascale Potholes for HPC: Execution Performance and Variability Analysis of the Flagship Application Code HemeLB
<p>Performance measurement and analysis of parallel applications is often challenging, despite many excellent commercial and open-source tools being available. Currently envisaged exascale computer systems exacerbate matters by requiring extremely high scalability to effectively exploit millions of processor cores. Unfortunately, significant application execution performance variability arising from increasingly complex interactions between hardware and system software makes this situation much more difficult for application developers and performance analysts alike.<br> </p> <p>This work considers the performance assessment of the HemeLB exascale flagship application code from the EU HPC Centre of Excellence (CoE) for Computational Biomedicine (CompBioMed) running on the SuperMUC-NG Tier-0 leadership system, using the methodology of the Performance Optimisation and Productivity (POP) CoE. Although 80% scaling efficiency is maintained to over 100,000 MPI processes, disappointing initial performance with more processes and corresponding poor strong scaling was identified to originate from the same few compute nodes in multiple runs, which later system diagnostic checks found had faulty DIMMs and lacklustre performance. Excluding these compute nodes from subsequent runs improved performance of executions with over 300,000 MPI processes by a factor of five, resulting in 190x speed-up compared to 864 MPI processes. While communication efficiency remains very good up to the largest scale, parallel efficiency is primarily limited by load balance found to be largely due to core-to-core and run-to-run variability from excessive stalls for memory accesses, that affect many HPC systems with Intel Xeon Scalable processors. The POP methodology for this performance diagnosis is demonstrated via a detailed exposition with widely deployed `standard' measurement and analysis tools.</p>
Cray-HPC Data Sample
<p>Sample of console and job logs of Cray system shared in the SC 18 <a href="https://dl.acm.org/doi/abs/10.5555/3291656.3291668">paper</a> (Doomsday: Predicting Which Node Will Fail When on Supercomputers). </p> <p>If you use this data, please cite the following paper that first released the logs:</p> <pre>@inproceedings{DBLP:conf/sc/DasMHRB18, author = {Anwesha Das and Frank Mueller and Paul Hargrove and Eric Roman and Scott B. Baden}, title = {Doomsday: predicting which node will fail when on supercomputers}, booktitle = {Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, {SC} 2018, Dallas, TX, USA, November 11-16}, pages = {9:1--9:14}, publisher = {{IEEE} / {ACM}}, year = {2018} } </pre>
HPC-GAP: Engineering a 21st-Century High-Performance Computer Algebra System
<p>Data for the experiments conducted in the paper</p>
Performance Measurement Dataset of the HPC Benchmarks FASTEST, Kripke, LULESH, MiniFE, Quicksilver, and RELeARN for Scalability Studies with Extra-P
<p>This dataset contains performance measurements of the HPC benchmarks FASTEST, Kripke, LULESH, MiniFE, Quicksilver, and RELeARN intended to be used for scalability studies with Extra-P (https://github.com/extra-p/extrap). The datasets contains measurements of various application configurations considering several model parameters, e.g., the number of MPI ranks and the input problem size, using weak scaling for each benchmark.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.