Topic modeling ensembles

  • Zhiyong Shen*
  • , Ping Luo
  • , Shengwen Yang
  • , Xukun Shen
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

In this paper we propose a framework of topic modeling ensembles, a novel solution to combine the models learned by topic modeling over each partition of the whole corpus. It has the potentials for applications such as distributed topic modeling for large corpora, and incremental topic modeling for rapidly growing corpora. Since only the base models, not the original documents, are required in the ensemble, all these applications can be performed in a privacy preserving manner. We explore the theoretical foundation of the proposed framework, give its geometric interpretation, and implement it for both PLSA and LDA. The evaluation of the implementations over the synthetic and real-life data sets shows that the proposed framework is much more efficient than modeling the original corpus directly while achieves comparable effectiveness in terms of perplexity and classification accuracy.

Original languageEnglish
JournalHP Laboratories Technical Report
Issue number158
StatePublished - 2010
Externally publishedYes

Keywords

  • Ensemble
  • Topic model

Fingerprint

Dive into the research topics of 'Topic modeling ensembles'. Together they form a unique fingerprint.

Cite this