Skip to main navigation Skip to search Skip to main content

Improving the shuffle of Hadoop MapReduce

  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

As an efficient parallel computing system based on MapReduce model, Hadoop is widely used for large-scale data analysis such as data mining, machine learning and scientific simulation. However, there are still some performance problems in MapReduce, especially the situation in the shuffle phase. In order to solve these problems, in this paper, a lightweight individual shuffle service component with more efficient I/O policy was proposed rather than the existing shuffle phase in MapReduce. We also describe how to implement the shuffle service in three steps: extract shuffle from reduce task as a shuffle task, reconstruct the shuffle task as a service and improve I/O scheduling policy on Map sides. Furthermore both simulated experiments and MapReduce job comparative studies are conducted to evaluate the performance of our improvements. The result reveals that our approach can decrease the whole job's execution time and make full use of cluster resources.

Original languageEnglish
Title of host publicationProceedings - IEEE 5th International Conference on Cloud Computing Technology and Science, CloudCom 2013
PublisherIEEE Computer Society
Pages266-273
Number of pages8
ISBN (Electronic)9780768550954
DOIs
StatePublished - 2013
Event5th IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2013 - Bristol, United Kingdom
Duration: 2 Dec 20135 Dec 2013

Publication series

NameProceedings of the International Conference on Cloud Computing Technology and Science, CloudCom
Volume1
ISSN (Print)2330-2194
ISSN (Electronic)2330-2186

Conference

Conference5th IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2013
Country/TerritoryUnited Kingdom
CityBristol
Period2/12/135/12/13

Keywords

  • Shuffle
  • hadoop
  • mapreduce

Fingerprint

Dive into the research topics of 'Improving the shuffle of Hadoop MapReduce'. Together they form a unique fingerprint.

Cite this