There is a newer version of the record available.

Published July 21, 2023 | Version 1.0

PM100: A Job Power Consumption Dataset of a Large-Scale HPC System

Description

The dataset is a collection of jobs extracted from the job_table data of M100 (https://doi.org/10.5281/zenodo.7588815), a collection of data extracted from a Tier-0 supercomputer hosted at CINECA (Marconi100, https://www.hpc.cineca.it/hardware/marconi100).  The original job data present in M100 are filtered out by considering only the jobs running exclusively on the resources. Each job entry included in PM100 contains the power consumption of the resources allocated to the job during its execution. The power consumption values are sampled every 20 seconds at node level, thus including the power contribution of all the components of the nodes (GPUs, CPUs, memory, etc...). The final dataset contains 231238 jobs, executed on Marconi100 between May and October 2020. 

The dataset is stored as a parquet file, where each entry contains the information on a job execution. 

The structure of the data, as well as the code to generate them, is contained in the official GitHub repository of the project: https://github.com/francescoantici/PM100-data/.

Files

Files (105.9 MB)

Name Size Download all
md5:2a692d9b9dc4a6ee4bc5e3eb3d0fa6f5
105.9 MB Download