There is a newer version of the record available.

Published November 13, 2020 | Version v1
Conference paper Open

Parallelized Data Replication of Multi-Petabyte Storage Systems

  • 1. DDN
  • 2. The University of Sydney

Contributors

Contact person:

Description

This paper presents the architecture of a highly parallelized data replication workflow implemented at The University of Sydney that forms the disaster recovery strategy for two 8-petabyte research data storage systems at the University. The solution leverages DDN’s GRIDScaler appliances, the information lifecycle management feature of the IBM Spectrum Scale File System, rsync, GNU Parallel and the MPI dsync tool from mpiFileUtils. It achieves high performance asynchronous data replication between two storage systems at sites 40km’s apart. In this paper, the methodology, performance benchmarks, technical challenges encountered and fine-tuning improvements in the implementation are presented.

Files

ws_hpcsysp103s1-file1.pdf

Files (886.7 kB)

Name Size Download all
md5:daf860846946d23154d65178924b6e2b
886.7 kB Preview Download