Published June 5, 2021 | Version v1

Dataset used to reproduce graphs in the SC21 submission "Understanding why machine learning models of I/O fail: A taxonomy of I/O throughput modelling errors"

Authors/Creators

  • 1. Texas A&M university

Description

Three datasets necessary to reproduce figures in our SC21 submission titled "Understanding why machine learning models of \\ I/O fail: A taxonomy of I/O throughput modelling errors". 

The darshan_theta_2017_2020.csv file is a CSV file constructed from Darshan logs, where every row represents an HPC job ran on ALCF Theta, and each column is a different feature of the job. This data is post-processed, in order to simplify reproduction of the paper. It is also anonymized, where the apps_short column represents the anonymized name of the application. 

The cobalt_theta_2017_2020.csv file contains Cobalt scheduler logs, where UIDs of allocations correspond to Darshan job UIDs. This data is also public, and is not preprocessed, only aggregated over 4 years.

The gauge_data.csv contains data from a single cluster of HPC jobs, collected using the Gauge tool (https://gauge.ascslab-tools.org). This data is a strict subset of the Darshan CSV listed above.

Files

cobalt_theta_2017_2020.csv

Files (1.3 GB)

Name Size
md5:2c97da3ebf29104e22b39edebd517281
224.7 MB Preview Download
md5:ed03bee6788165a180fa5c48b4cede99
1.1 GB Preview Download
md5:d9b3fbbcae8601a274064813845b4ad3
16.9 kB Preview Download