跳到主要导航 跳到搜索 跳到主要内容

Theme extraction from chinese web documents based on page segmentation and entropy

  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Web pages often contain "clutters" (defined by us as unnecessary images, navigational menus and extraneous Ad links) around the body of an article that may distract users from the actual content. Therefore, how to extract useful and relevant themes from these web pages becomes a research focus. This paper proposes a new method for web theme extraction. The method firstly uses page segmentation technique to divide a web page into many unrelated blocks, and then calculates entropy of each block and that of the entire web page, then prunes redundant blocks whose entropies are larger than the threshold of the web page, lastly exports the rest blocks as theme of the web page. Moreover, it is verified by experiments that the new method takes better effect on theme extraction from Chinese web pages.

源语言英语
主期刊名Foundations of Intelligent Systems - 18th International Symposium, ISMIS 2009, Proceedings
221-230
页数10
DOI
出版状态已出版 - 2009
活动18th International Symposium on Methodologies for Intelligent Systems, ISMIS 2009 - Prague, 捷克共和国
期限: 14 9月 200917 9月 2009

出版系列

姓名Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
5722 LNAI
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议18th International Symposium on Methodologies for Intelligent Systems, ISMIS 2009
国家/地区捷克共和国
Prague
时期14/09/0917/09/09

学术指纹

探究 'Theme extraction from chinese web documents based on page segmentation and entropy' 的科研主题。它们共同构成独一无二的学术指纹。

引用此