跳到主要导航 跳到搜索 跳到主要内容

Block-level links based content extraction

  • Shixing Shen*
  • , Hui Zhang
  • *此作品的通讯作者
  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

We present block-level links based content extraction (BLCE) - a method to extract content from the web pages by using the link attributes of blocks, which contains the number of links and the length of link text (anchor text). We describe how to divide one webpage into blocks and how to merge the similar blocks into one, then compute the number of links and the total length of anchor text. We find that extracting content only with the number of links and length of anchor text is not effective because the number of links and length of link text are proportional to the length of page. Density of links is a good method to solve this. So we use the content links ratios and the content anchor text ratios to describe the link attribute of the blocks. BLCE performs better than other methods especially in the new web pages with DIV and CSS where traditional algorithm can't work well.

源语言英语
主期刊名Proceedings - 2011 4th International Symposium on Parallel Architectures, Algorithms and Programming, PAAP 2011
330-333
页数4
DOI
出版状态已出版 - 2011
活动2011 4th International Symposium on Parallel Architectures, Algorithms and Programming, PAAP 2011 - Tianjin, 中国
期限: 9 12月 201111 12月 2011

出版系列

姓名Proceedings - 2011 4th International Symposium on Parallel Architectures, Algorithms and Programming, PAAP 2011

会议

会议2011 4th International Symposium on Parallel Architectures, Algorithms and Programming, PAAP 2011
国家/地区中国
Tianjin
时期9/12/1111/12/11

学术指纹

探究 'Block-level links based content extraction' 的科研主题。它们共同构成独一无二的学术指纹。

引用此