跳到主要导航 跳到搜索 跳到主要内容

A CRF-based approach for web object extraction

  • Rui Liu*
  • , Rui Xiong
  • , Kun Gao
  • *此作品的通讯作者
  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

A method for extracting Web object is presented in this paper. Firstly, Web object blocks are obtained by blocking the web page and calculating the information entropy of it. Then it uses Conditional Random Field model as a probability and statistics model, and builds a series of feature templates according to the characteristics of objects themselves. Feature functions are generated based on the result of Chinese word segmentation and feature templates. It uses a limited memory BFGS algorithm to estimate parameters of the model, and labels property sequences of Web object blocks by Viterbi algorithm. Experiment result shows that the proposed method is an effective way to extract science data.

源语言英语
主期刊名Proceedings - 2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010
483-487
页数5
DOI
出版状态已出版 - 2010
活动2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010 - Chengdu, 中国
期限: 9 7月 201011 7月 2010

出版系列

姓名Proceedings - 2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010
4

会议

会议2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010
国家/地区中国
Chengdu
时期9/07/1011/07/10

学术指纹

探究 'A CRF-based approach for web object extraction' 的科研主题。它们共同构成独一无二的学术指纹。

引用此