Published June 18, 2023 | Version v1
Journal article Open

BtrBlocks: Efficient Columnar Compression for Data Lakes

  • 1. Technische Universität München
  • 2. Friedrich-Alexander-Universität Erlangen-Nürnberg

Description

Analytics is moving to the cloud and data is moving into data lakes. These reside on object storage services like S3 and enable seamless data sharing and system interoperability. To support this, many systems build on open storage formats like Apache Parquet. However, these formats are not optimized for remotely-accessed data lakes and today's high-throughput networks. Inefficient decompression makes scans CPU-bound and thus increases query time and cost. With this work we present BtrBlocks, an open columnar storage format designed for data lakes. BtrBlocks uses a set of lightweight encoding schemes, achieving fast and efficient decompression and high compression ratios.

 

Files

btrblocks.pdf

Files (840.6 kB)

Name Size Download all
md5:f14a190ffa043fc9501cf8d7e78f7c76
840.2 kB Preview Download
md5:9bf9732e43df0a668dec342ae4198a55
321 Bytes Download

Additional details

Funding

CODAC – Commoditizing Data Analytics in the Cloud 101041375
European Commission