TY - JOUR
T1 - Column-wise compression of open relational data
AU - Wandelt, Sebastian
AU - Sun, Xiaoqian
AU - Leser, Ulf
N1 - Publisher Copyright:
© 2018 Elsevier Inc.
PY - 2018/8
Y1 - 2018/8
N2 - The recent growth of open data initiatives has led to a tremendous increase in publicly available data resources. This amount of data, together with a rising interest in people to analyze it, poses severe challenges regarding data storage. Data suppliers often compress their resources with standard compressors. The choice of a compression technique, however, has significant impacts on the compression ratio, the compression speed, and the decompression speed. In this paper, we provide an empirical analysis on the compression of open data provided in a relational format, such as comma-separated value files. We consider several compression tools and parameter settings. Furthermore, we propose using a novel column-wise compression strategy, where items that have similar properties, are compressed together. We perform a comprehensive analysis on 24 datasets from different domains, such as life sciences, governmental data, finance sector, and public transportation, which cover a wide range of file sizes (from a few MB to several GB). Our results show that the traversal strategy is of paramount importance for achieving high compression ratios; with improvements of up to one order of magnitude. This study further highlights a set of issues for future work on compressing open data.
AB - The recent growth of open data initiatives has led to a tremendous increase in publicly available data resources. This amount of data, together with a rising interest in people to analyze it, poses severe challenges regarding data storage. Data suppliers often compress their resources with standard compressors. The choice of a compression technique, however, has significant impacts on the compression ratio, the compression speed, and the decompression speed. In this paper, we provide an empirical analysis on the compression of open data provided in a relational format, such as comma-separated value files. We consider several compression tools and parameter settings. Furthermore, we propose using a novel column-wise compression strategy, where items that have similar properties, are compressed together. We perform a comprehensive analysis on 24 datasets from different domains, such as life sciences, governmental data, finance sector, and public transportation, which cover a wide range of file sizes (from a few MB to several GB). Our results show that the traversal strategy is of paramount importance for achieving high compression ratios; with improvements of up to one order of magnitude. This study further highlights a set of issues for future work on compressing open data.
UR - https://www.scopus.com/pages/publications/85047090790
U2 - 10.1016/j.ins.2018.04.074
DO - 10.1016/j.ins.2018.04.074
M3 - 文章
AN - SCOPUS:85047090790
SN - 0020-0255
VL - 457-458
SP - 48
EP - 61
JO - Information Sciences
JF - Information Sciences
ER -