跳到主要导航 跳到搜索 跳到主要内容

Exploiting associations between word clusters and document classes for cross-domain text categorization

  • Fuzhen Zhuang*
  • , Ping Luo
  • , Hui Xiong
  • , Qing He
  • , Yuhong Xiong
  • , Zhongzhi Shi
  • *此作品的通讯作者
  • CAS - Institute of Computing Technology
  • University of Chinese Academy of Sciences
  • Hewlett-Packard
  • Rutgers University
  • Innovation Works

科研成果: 期刊稿件文章同行评审

摘要

Cross-domain text categorization targets on adapting the knowledge learnt from a labeled source domain to an unlabeled target domain, where the documents from the source and target domains are drawn from different distributions. However, in spite of the different distributions in raw-word features, the associations between word clusters (conceptual features) and document classes may remain stable across different domains. In this paper, we exploit these unchanged associations as the bridge of knowledge transformation from the source domain to the target domain by the non-negative matrix tri-factorization. Specifically, we formulate a joint optimization framework of the two matrix tri-factorizations for the source- and target-domain data, respectively, in which the associations between word clusters and document classes are shared between them. Then, we give an iterative algorithm for this optimization and theoretically show its convergence. The comprehensive experiments show the effectiveness of this method. In particular, we show that the proposed method can deal with some difficult scenarios where baseline methods usually do not perform well.

源语言英语
页(从-至)100-114
页数15
期刊Statistical Analysis and Data Mining
4
1
DOI
出版状态已出版 - 2月 2011
已对外发布

学术指纹

探究 'Exploiting associations between word clusters and document classes for cross-domain text categorization' 的科研主题。它们共同构成独一无二的学术指纹。

引用此