Performance optimization of legacy HDF5 data for Web Object Stores
Description
In this poster we examine the relative benefits of storing data in S3 using HDF5 and Zarr in the context of NASA's Earthdata cloud computing environment. We are particularly interested in supporting users who do not have prior experience with optimizing data for cloud computing systems. Many NASA datasets are stored in HDF5 files and using these files without reformatting the data would be the simplest path to use. However, HDF5 was designed long before Web Object Store (WOS) systems were developed and its API realizations have typically depended on the ability of reader software to perform rapid low-latency random access I/O operations. Zarr is a new format/API that shards data in separately addressable objects, making access to portions of data more straightforward with WOS systems. Existing studies have shown that the performance of the Zarr and HDF5 APIs are dependent on a complex stack of I/O middleware and that performance optimization is non-trivial (Kang, D., Rübel, O., Byna, S. and Blanas, S., 2019). In some cases, Zarr outperforms HDF5 and in others, the reverse is true (Pfander, I., Johnson, H. and Arms, S., 2021). If, for some use cases, data stored in HDF5 can be efficiently accessed without first reformatting them in Zarr, users can be saved the effort to learn how to optimize data organization. We extend the performance analysis of HDF5 and Zarr in the context of WOS to include the DMR++ technique OPeNDAP has developed to access HDF5 data. In particular, we examine the issue of 'chunk size' and how that affects the data access efficiency of WOS systems. We examine technology that treats a series of small and contiguous chunks in an HDF5 file as one 'super chunk' and examine the effect of that optimization on WOS data access performance.
Files
Gallagher_AGU_2022_Poster.pdf
Files
(6.9 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:f6c8b8d3a0abe1adb650547fd3214b6b
|
6.9 MB | Preview Download |