Published July 30, 2013 | Version v1

Extending SLURM with Support for GPU Ranges

Authors/Creators

  • 1. Computer Engineering Department, Bogazici University, Istanbul, Turkey

Contributors

  • 1. Computer Engineering Department, Bogazici University, Istanbul, Turkey

Description

SLURM resource management system is used on many TOP500 supercomputers. In this work, we present enhancements
that we added to our AUCSCHED heterogeneous CPU-GPU scheduler plug-in whose rst version was released in
December 2012. In this new version, called AUCSCHED2, two enhancements are contributed: The rst is the extension
of SLURM to support GPU ranges. The current version of SLURM supports speci cation of node range but not of GPU
ranges. Such a feature can be very useful to runtime auto-tuning applications and systems that can make use of variable
number of GPUs. The second enhancement involves the implementation of a new integer programming formulation in
AUCSCHED2 that drastically reduces the number of variables. This allows faster solution and larger number of bids to
be generated. SLURM emulation results are presented for the heterogeneous 1408 node Tsubame supercomputer which
has 12 cores and 3 GPU's on each of its nodes. AUCSCHED2 is available at http://code.google.com/p/slurm-ipsched/.

Files

WP80.pdf

Files (297.0 kB)

Name Size Download all
md5:4057db9d41b3de63f50663e50608ee98
297.0 kB Preview Download

Additional details

Funding

European Commission
PRACE-2IP - PRACE - Second Implementation Phase Project 283493