Skip to main navigation Skip to search Skip to main content

K-means method for grouping in hybrid mapreduce clusters

  • Yang Yang
  • , Xiang Long
  • , Bo Jiang*
  • , Yu Liu
  • *Corresponding author for this work
  • Beihang University

Research output: Contribution to journalArticlepeer-review

Abstract

In hybrid cloud computing era, hybrid clusters which are consisted of virtual machines and physical machines become more and more Popular?. MapReduce is a good weapon in this big data era where social computing and multimedia computing are emerging. One of the biggest challenges in hybrid mapreduce cluster is I/O bottleneck which would be aggravated under big data computing. In this paper, we take data locality into consideration and group slave nodes with low intra-communication and high intracommunication. After introducing the architecture and implementation of our grouped hybrid mapreduce cluster (GHMC), we give our k-means algorithm in GHMC and evaluate it with reality environments. The results show that there is a nearly 34.9% performance improvement in our system achieved by the K-means algorithm. Moreover, GHMC system also shows good scalability.

Original languageEnglish
Pages (from-to)383-388
Number of pages6
JournalJournal of Theoretical and Applied Information Technology
Volume48
Issue number1
StatePublished - 2013

Keywords

  • Hybrid cloud computing
  • K-means cluster
  • MapReduce

Fingerprint

Dive into the research topics of 'K-means method for grouping in hybrid mapreduce clusters'. Together they form a unique fingerprint.

Cite this