Accelerating SSSP for Power-Law Graphs
Description
The single-source shortest path (SSSP) problem is one of the most important and well-studied graph problems widely used in many application domains, such as road navigation, neural image reconstruction, and social network analysis. Although we have known various SSSP algorithms for decades, implementing one for large-scale power-law graphs efficiently is still highly challenging today, because ① a work-efficient SSSP algorithm requires priority-order traversal of graph data, ② the priority queue needs to be scalable both in throughput and capacity, and ③ priority-order traversal requires extensive random memory accesses on graph data.
In this paper, we present SPLAG to accelerate SSSP for power-law graphs on FPGAs. SPLAG uses a coarse-grained priority queue (CGPQ) to enable high-throughput priority-order graph traversal with a large frontier. To mitigate the high-volume random accesses, SPLAG employs a customized vertex cache (CVC) to reduce off-chip memory access and improve the throughput to read and update vertex data. Experimental results on various synthetic and real-world datasets show up to a 4.9× speedup over state-of-the-art SSSP accelerators, a 2.6× speedup over 32-thread CPU running at 4.4 GHz, and a 0.9× speedup over an A100 GPU that has 4.1× power budget and 3.4× HBM bandwidth. Such a high performance would place SPLAG in the 14th position of the Graph 500 benchmark for data intensive applications (the highest using a single FPGA) with only a 45 W power budget. SPLAG is written in high-level synthesis C++ and is fully parameterized, which means it can be easily ported to various different FPGAs with different configurations. SPLAG is open-source at https://github.com/UCLA-VAST/splag.
To completely reproduce the experiments using the pre-built docker images, the followings are required:
- Docker on a x64 Linux server. To run the GPU experiments, the Nvidia container toolkit is required additionally.
- Xilinx Alveo U280 FPGA with the
xilinx_u280_xdma_201920_3platform. This is the FPGA used in most experiments. - Xilnix Alveo U250 FPGA with the
xilinx_u250_xdma_201830_2platform. Without this FPGA, you won't be able to reproduce the comparison with ThunderGP and HitGraph. - Dual-socket Intel Xeon Gold 6244 CPU. Using a different CPU may lead to a very different comparison between SPLAG on U280 and CPU.
- Nvidia A100 40 GB GPU. Without this GPU, you won't be able to reproduce the comparison between SPLAG on U280 and GPU.
- XRT
2.11.634as shown byxbutil version. Using a different version of XRT will build the docker image from source. - CUDA 11.3 as shown by
nvidia-smi. Using a different version of CUDA will build the docker image from source. - Vitis HLS 2020.2 and Vitis 2021.1, if you would like to build the FPGA bitstreams from source. Using different versions may lead to very different quality of results. Pre-built bitstreams are available and will be used by default.
The first step to reproduce the experiments is to download all the files in the same directory. After that, create the data directory by extracting data.tgz as follows. It is ~10 GB large after extraction.
tar xvf data.tgz
The data directory contains three pre-built FPGA bitstreams (and many other stuff). If you would like to build the bitstreams from source, you can do as follows. You should modify the two environment variables properly according to your local setup. The three bitstreams are for U280, VU5P, and U250, respectively. You may run them in parallel if you have sufficient memory. Generating a bitstream takes ~14 h.
XILINX_HLS=/opt/tools/xilinx/Vitis_HLS/2020.2 XILINX_VIVADO=/opt/tools/xilinx/Vivado/2021.1 ./run.sh build-u280
cp build/u280/SSSP.xilinx_u280_xdma_201920_3.hw.xclbin data/SSSP.xilinx_u280_xdma_201920_3.hw.xclbin
XILINX_HLS=/opt/tools/xilinx/Vitis_HLS/2020.2 XILINX_VIVADO=/opt/tools/xilinx/Vivado/2021.1 ./run.sh build-vu5p
cp build/vu5p/SSSP.xilinx_u250_xdma_201830_2.hw.xclbin data/SSSP.xilinx_u250_xdma_201830_2.vu5p.hw.xclbin
XILINX_HLS=/opt/tools/xilinx/Vitis_HLS/2020.2 XILINX_VIVADO=/opt/tools/xilinx/Vivado/2021.1 ./run.sh build-u250
cp build/u250/SSSP.xilinx_u250_xdma_201830_2.hw.xclbin data/SSSP.xilinx_u250_xdma_201830_2.hw.xclbin
To run the FPGA experiments, do as follows. It will take ~1 h.
./run.sh fpga
To run the CPU experiments, do as follows. Note that the output directory generated from the FPGA experiments are required. If you would like to run the CPU expriments on a different machine, make sure to copy the output directory.
./run.sh cpu
To run the GPU experiments, do as follows. Similar to the CPU experimetns, the output directory generated from the FPGA experiments are required. If you would like to run the GPU expriments on a different machine, make sure to copy the output directory.
./run.sh gpu
Files
Files
(6.7 GB)
| Name | Size | |
|---|---|---|
|
md5:07b30c0043516169f4794add30a5bc9e
|
5.3 GB | Download |
|
md5:1fb6bec404824dc251abdc4e8f15b0f0
|
5.0 kB | Download |
|
md5:7fb96d144ae322160ccc4ae7f399993e
|
4.7 kB | Download |
|
md5:dbdd74249dd0fa386a53d194dd59d111
|
46.5 MB | Download |
|
md5:f6054d323f87bff31088ce25f6ae41e8
|
345.8 MB | Download |
|
md5:e73e6d804690dd5776f8a9a12c262610
|
1.0 GB | Download |