There is a newer version of the record available.

Published February 27, 2022 | Version v1

Accelerating SSSP for Power-Law Graphs

  • 1. UCLA

Description

The single-source shortest path (SSSP) problem is one of the most important and well-studied graph problems widely used in many application domains, such as road navigation, neural image reconstruction, and social network analysis. Although we have known various SSSP algorithms for decades, implementing one for large-scale power-law graphs efficiently is still highly challenging today, because ① a work-efficient SSSP algorithm requires priority-order traversal of graph data, ② the priority queue needs to be scalable both in throughput and capacity, and ③ priority-order traversal requires extensive random memory accesses on graph data.

In this paper, we present SPLAG to accelerate SSSP for power-law graphs on FPGAs. SPLAG uses a coarse-grained priority queue (CGPQ) to enable high-throughput priority-order graph traversal with a large frontier. To mitigate the high-volume random accesses, SPLAG employs a customized vertex cache (CVC) to reduce off-chip memory access and improve the throughput to read and update vertex data. Experimental results on various synthetic and real-world datasets show up to a 4.9× speedup over state-of-the-art SSSP accelerators, a 2.6× speedup over 32-thread CPU running at 4.4 GHz, and a 0.9× speedup over an A100 GPU that has 4.1× power budget and 3.4× HBM bandwidth. Such a high performance would place SPLAG in the 14th position of the Graph 500 benchmark for data intensive applications (the highest using a single FPGA) with only a 45 W power budget. SPLAG is written in high-level synthesis C++ and is fully parameterized, which means it can be easily ported to various different FPGAs with different configurations. SPLAG is open-source at https://github.com/UCLA-VAST/splag.

To completely reproduce the experiments using the pre-built docker images, the followings are required:

  • Docker on a x64 Linux server. To run the GPU experiments, the Nvidia container toolkit is required additionally.
  • Xilinx Alveo U280 FPGA with the xilinx_u280_xdma_201920_3 platform. This is the FPGA used in most experiments.
  • Xilnix Alveo U250 FPGA with the xilinx_u250_xdma_201830_2 platform. Without this FPGA, you won't be able to reproduce the comparison with ThunderGP and HitGraph.
  • Dual-socket Intel Xeon Gold 6244 CPU. Using a different CPU may lead to a very different comparison between SPLAG on U280 and CPU.
  • Nvidia A100 40 GB GPU. Without this GPU, you won't be able to reproduce the comparison between SPLAG on U280 and GPU.
  • XRT 2.11.634 as shown by xbutil version. Using a different version of XRT will build the docker image from source.
  • CUDA 11.3 as shown by nvidia-smi. Using a different version of CUDA will build the docker image from source.
  • Vitis HLS 2020.2 and Vitis 2021.1, if you would like to build the FPGA bitstreams from source. Using different versions may lead to very different quality of results. Pre-built bitstreams are available and will be used by default.

The first step to reproduce the experiments is to download all the files in the same directory. After that, create the data directory by extracting data.tgz as follows. It is ~10 GB large after extraction.

tar xvf data.tgz

The data directory contains three pre-built FPGA bitstreams (and many other stuff). If you would like to build the bitstreams from source, you can do as follows. You should modify the two environment variables properly according to your local setup. The three bitstreams are for U280, VU5P, and U250, respectively. You may run them in parallel if you have sufficient memory. Generating a bitstream takes ~14 h.

XILINX_HLS=/opt/tools/xilinx/Vitis_HLS/2020.2 XILINX_VIVADO=/opt/tools/xilinx/Vivado/2021.1 ./run.sh build-u280
cp build/u280/SSSP.xilinx_u280_xdma_201920_3.hw.xclbin data/SSSP.xilinx_u280_xdma_201920_3.hw.xclbin
XILINX_HLS=/opt/tools/xilinx/Vitis_HLS/2020.2 XILINX_VIVADO=/opt/tools/xilinx/Vivado/2021.1 ./run.sh build-vu5p
cp build/vu5p/SSSP.xilinx_u250_xdma_201830_2.hw.xclbin data/SSSP.xilinx_u250_xdma_201830_2.vu5p.hw.xclbin
XILINX_HLS=/opt/tools/xilinx/Vitis_HLS/2020.2 XILINX_VIVADO=/opt/tools/xilinx/Vivado/2021.1 ./run.sh build-u250
cp build/u250/SSSP.xilinx_u250_xdma_201830_2.hw.xclbin data/SSSP.xilinx_u250_xdma_201830_2.hw.xclbin

To run the FPGA experiments, do as follows. It will take ~1 h.

./run.sh fpga

To run the CPU experiments, do as follows. Note that the output directory generated from the FPGA experiments are required. If you would like to run the CPU expriments on a different machine, make sure to copy the output directory.

./run.sh cpu

To run the GPU experiments, do as follows. Similar to the CPU experimetns, the output directory generated from the FPGA experiments are required. If you would like to run the GPU expriments on a different machine, make sure to copy the output directory.

./run.sh gpu

Files

Files (6.7 GB)

Name Size
md5:07b30c0043516169f4794add30a5bc9e
5.3 GB Download
md5:1fb6bec404824dc251abdc4e8f15b0f0
5.0 kB Download
md5:7fb96d144ae322160ccc4ae7f399993e
4.7 kB Download
md5:dbdd74249dd0fa386a53d194dd59d111
46.5 MB Download
md5:f6054d323f87bff31088ce25f6ae41e8
345.8 MB Download
md5:e73e6d804690dd5776f8a9a12c262610
1.0 GB Download