Skip to main navigation Skip to search Skip to main content

Improving MapReduce performance by using a new partitioner in YARN

  • Wei Lu
  • , Lei Chen*
  • , Haitao Yuan
  • , Weiwei Xing
  • , Liqiang Wang
  • , Yong Yang
  • *Corresponding author for this work
  • Beijing Jiaotong University
  • University of Central Florida

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Data skew, cluster heterogeneity, and network traffic are three issues that significantly influence the performance of MapReduce applications. However, the Hash-Partitioner in native Hadoop does not consider them. This paper proposes a new partitioner in Yarn (Hadoop 2.6.0), namely, PIY, which adopts an innovative parallel sampling method to achieve the distribution of the intermediate data. Based on this, firstly, PIY mitigates data skew in MapReduce applications. Secondly, PIY considers the heterogeneity of the computing resource to balance the load among Reducers. Thirdly, PIY reduces the network traffic in shuffle phase by trying to retain intermediate data on those nodes who act as both mapper and reducer. Compared with the native Hadoop and some other popular strategies, PIY can reduce the execution time by 35.62% and 50.65% in homogeneous and heterogeneous cluster, respectively. We also implement PIY in parallel image processing. Compared with several existing strategies, PIY can reduce the execution time by 11.2%.

Original languageEnglish
Title of host publicationProceedings - DMSVLSS 2017
Subtitle of host publication23rd International Conference on Distributed Multimedia Systems, Visual Languages and Sentient Systems
PublisherKnowledge Systems Institute Graduate School
Pages24-33
Number of pages10
ISBN (Electronic)189170642X, 9781891706424
DOIs
StatePublished - 2017
Externally publishedYes
Event23rd International Conference on Distributed Multimedia Systems, Visual Languages and Sentient Systems, DMSVLSS 2017 - Pittsburgh, United States
Duration: 7 Jul 20178 Jul 2017

Publication series

NameProceedings - DMSVLSS 2017: 23rd International Conference on Distributed Multimedia Systems, Visual Languages and Sentient Systems

Conference

Conference23rd International Conference on Distributed Multimedia Systems, Visual Languages and Sentient Systems, DMSVLSS 2017
Country/TerritoryUnited States
CityPittsburgh
Period7/07/178/07/17

Keywords

  • Data skew
  • Data transmission amount
  • Hadoop
  • Heterogeneousparallel image processing
  • Load balance
  • MapReduce

Fingerprint

Dive into the research topics of 'Improving MapReduce performance by using a new partitioner in YARN'. Together they form a unique fingerprint.

Cite this