跳到主要导航 跳到搜索 跳到主要内容

Research on mixture language model-based document clustering

  • National University of Defense Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Language modeling with semantic smoothing is proposed as an effective way to improve the quality of document clustering. However, the existing semantic smoothing model is not effective for partitional clustering because it can not assign fit weight to "general" word in a collection. In this paper, inspired by mixture probability model, we put forward a mixture language model for document clustering. The new model can alleviate the effect of "general" word, simultaneously, it can integrate the context information and solve the polysemy problems in a document. Based the new model, an EM algorithm for partitional clustering is present. The experimental results show our algorithms are more effective than the previous methods to improve the cluster quality.

源语言英语
主期刊名2008 IEEE International Conference on Granular Computing, GRC 2008
649-652
页数4
DOI
出版状态已出版 - 2008
活动2008 IEEE International Conference on Granular Computing, GRC 2008 - Hangzhou, 中国
期限: 26 8月 200828 8月 2008

出版系列

姓名2008 IEEE International Conference on Granular Computing, GRC 2008

会议

会议2008 IEEE International Conference on Granular Computing, GRC 2008
国家/地区中国
Hangzhou
时期26/08/0828/08/08

指纹

探究 'Research on mixture language model-based document clustering' 的科研主题。它们共同构成独一无二的指纹。

引用此