TY - GEN
T1 - Improving MapReduce performance by using a new partitioner in YARN
AU - Lu, Wei
AU - Chen, Lei
AU - Yuan, Haitao
AU - Xing, Weiwei
AU - Wang, Liqiang
AU - Yang, Yong
PY - 2017
Y1 - 2017
N2 - Data skew, cluster heterogeneity, and network traffic are three issues that significantly influence the performance of MapReduce applications. However, the Hash-Partitioner in native Hadoop does not consider them. This paper proposes a new partitioner in Yarn (Hadoop 2.6.0), namely, PIY, which adopts an innovative parallel sampling method to achieve the distribution of the intermediate data. Based on this, firstly, PIY mitigates data skew in MapReduce applications. Secondly, PIY considers the heterogeneity of the computing resource to balance the load among Reducers. Thirdly, PIY reduces the network traffic in shuffle phase by trying to retain intermediate data on those nodes who act as both mapper and reducer. Compared with the native Hadoop and some other popular strategies, PIY can reduce the execution time by 35.62% and 50.65% in homogeneous and heterogeneous cluster, respectively. We also implement PIY in parallel image processing. Compared with several existing strategies, PIY can reduce the execution time by 11.2%.
AB - Data skew, cluster heterogeneity, and network traffic are three issues that significantly influence the performance of MapReduce applications. However, the Hash-Partitioner in native Hadoop does not consider them. This paper proposes a new partitioner in Yarn (Hadoop 2.6.0), namely, PIY, which adopts an innovative parallel sampling method to achieve the distribution of the intermediate data. Based on this, firstly, PIY mitigates data skew in MapReduce applications. Secondly, PIY considers the heterogeneity of the computing resource to balance the load among Reducers. Thirdly, PIY reduces the network traffic in shuffle phase by trying to retain intermediate data on those nodes who act as both mapper and reducer. Compared with the native Hadoop and some other popular strategies, PIY can reduce the execution time by 35.62% and 50.65% in homogeneous and heterogeneous cluster, respectively. We also implement PIY in parallel image processing. Compared with several existing strategies, PIY can reduce the execution time by 11.2%.
KW - Data skew
KW - Data transmission amount
KW - Hadoop
KW - Heterogeneousparallel image processing
KW - Load balance
KW - MapReduce
UR - https://www.scopus.com/pages/publications/85029592551
U2 - 10.18293/DMSVLSS2017-002
DO - 10.18293/DMSVLSS2017-002
M3 - 会议稿件
AN - SCOPUS:85029592551
T3 - Proceedings - DMSVLSS 2017: 23rd International Conference on Distributed Multimedia Systems, Visual Languages and Sentient Systems
SP - 24
EP - 33
BT - Proceedings - DMSVLSS 2017
PB - Knowledge Systems Institute Graduate School
T2 - 23rd International Conference on Distributed Multimedia Systems, Visual Languages and Sentient Systems, DMSVLSS 2017
Y2 - 7 July 2017 through 8 July 2017
ER -