TY - GEN
T1 - Accelerating Big Data Application by Eliminating Redundancy on Hadoop Cluster
AU - Lei, Kelun
AU - Du, Shaokang
AU - You, Xin
AU - Xuan, Zhibo
AU - Kong, Haoran
AU - Yang, Hailong
AU - Shang, Jing
AU - Xiao, Zhiwen
AU - Wu, Zhihui
AU - Luan, Zhongzhi
AU - Qian, Depei
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Big data applications are widely adopted to mine valuable information from a tremendous amount of industry data, which is commonly represented as a series of map-reduce operations. Among various map-reduce frameworks, Hadoop is most commonly adopted for data processing at large scale. Although Hadoop eases the development of highly scalable distributed big data applications, inefficient implementation due to poor coding practice and deep software abstractions can cause severe performance issues such as unreasonable slowdown, high response latency, and waste of computing resources, which can lead to unsatisfactory serving delay or significant maintenance cost. In this paper, we first categorize three common types of redundant patterns in big data applications. Then we propose a tool-assisted optimization workflow to detect the redundant patterns automatically, which profiles the application by sampling hardware performance monitoring units. Moreover, we present a profiling visualization method that can help to pinpoint the redundant codes. Based on these approaches, we optimize several big data applications by eliminating redundancies, yielding up to 14.8% performance improvement.
AB - Big data applications are widely adopted to mine valuable information from a tremendous amount of industry data, which is commonly represented as a series of map-reduce operations. Among various map-reduce frameworks, Hadoop is most commonly adopted for data processing at large scale. Although Hadoop eases the development of highly scalable distributed big data applications, inefficient implementation due to poor coding practice and deep software abstractions can cause severe performance issues such as unreasonable slowdown, high response latency, and waste of computing resources, which can lead to unsatisfactory serving delay or significant maintenance cost. In this paper, we first categorize three common types of redundant patterns in big data applications. Then we propose a tool-assisted optimization workflow to detect the redundant patterns automatically, which profiles the application by sampling hardware performance monitoring units. Moreover, we present a profiling visualization method that can help to pinpoint the redundant codes. Based on these approaches, we optimize several big data applications by eliminating redundancies, yielding up to 14.8% performance improvement.
KW - Hadoop
KW - Performance Optimization
KW - Redundant Patterns
UR - https://www.scopus.com/pages/publications/85190254708
U2 - 10.1109/ICPADS60453.2023.00114
DO - 10.1109/ICPADS60453.2023.00114
M3 - 会议稿件
AN - SCOPUS:85190254708
T3 - Proceedings of the International Conference on Parallel and Distributed Systems - ICPADS
SP - 751
EP - 756
BT - Proceedings - 2023 IEEE 29th International Conference on Parallel and Distributed Systems, ICPADS 2023
PB - IEEE Computer Society
T2 - 29th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2023
Y2 - 17 December 2023 through 21 December 2023
ER -