Fives Input dataset (Cobalt & Darshan traces, combined and preprocessed)
<p>Dataset made of aggregated and curated Cobalt and Darshan logs from the Theta HPC platform at ALCF.</p> <p>Cobalt and Darshan logs were obtained from ALCF Public Data repository (https://reports.alcf.anl.gov/data/index.html) and cover the year 2022. This data was generated from resources of the Argonne Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC02-06CH11357. In order to use the scripts contained within this archive, these datasets must be downloaded and placed in the directory '2022' at the root of the extracted archive.</p> <p>The Darshan logs used in this datasets are originillay available in an aggregated form. The levels of details are usually the following : </p> <ul> <li>job (reservation made to a resource manager for some platform resources)</li> <li>application run (application running inside the job, on the reserved resources ; there may be multiple ones, sequentially or in parallel, during a job's execution)</li> <li>I/O operation (read or write registered to a file from a process of an application)</li> </ul> <p>Darshan CSV files for Theta contain job and application runs informations, but individual I/O of each application run is aggregated into a single entry.</p> <p>This resource is organised as a single archive containing:</p> <ul> <li>YAML files with our datasets, at various granularity levels (in 'preprocessed_datastets' directory): <ul> <li>48 files containing each<strong> 1 month worth of job traces</strong> for one of <strong>3 job classes</strong> (4 files per month, one per job class and one with all job classes) </li> <li>4 files containing each the entire year worth of job traces ; 1 file per job class, 1 file with all job classes.</li> </ul> </li> <li>A Jupyter Lab notebook, which contains the necessary routines to create aformentionned datasets from raw logs files from ALCF, for the Theta system</li> <li>A requirements.txt file, describing required Python packages and their versions.</li> <li>Various empty directories meant to receive outputs from the Jupyter notebook.</li> </ul>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4