Skip to main navigation Skip to search Skip to main content

Validation of overlapping clustering: A random clustering perspective

  • Junjie Wu*
  • , Hua Yuan
  • , Hui Xiong
  • , Guoqing Chen
  • *Corresponding author for this work
  • University of Electronic Science and Technology of China
  • Rutgers - The State University of New Jersey, Newark
  • Tsinghua University

Research output: Contribution to journalArticlepeer-review

Abstract

As a widely used clustering validation measure, the F-measure has received increased attention in the field of information retrieval. In this paper, we reveal that the F-measure can lead to biased views as to results of overlapped clusters when it is used for validating the data with different cluster numbers (incremental effect) or different prior probabilities of relevant documents (prior-probability effect). We propose a new "IMplication Intensity" (IMI) measure which is based on the F-measure and is developed from a random clustering perspective. In addition, we carefully investigate the properties of IMI. Finally, experimental results on real-world data sets show that IMI significantly alleviates biased incremental and prior-probability effects which are inherent to the F-measure.

Original languageEnglish
Pages (from-to)4353-4369
Number of pages17
JournalInformation Sciences
Volume180
Issue number22
DOIs
StatePublished - 15 Nov 2010

Keywords

  • Cluster validation
  • F-measure
  • Implication intensity (IMI)
  • Incomplete beta function
  • Information retrieval

Fingerprint

Dive into the research topics of 'Validation of overlapping clustering: A random clustering perspective'. Together they form a unique fingerprint.

Cite this