Skip to main navigation Skip to search Skip to main content

A CRF-based approach for web object extraction

  • Rui Liu*
  • , Rui Xiong
  • , Kun Gao
  • *Corresponding author for this work
  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

A method for extracting Web object is presented in this paper. Firstly, Web object blocks are obtained by blocking the web page and calculating the information entropy of it. Then it uses Conditional Random Field model as a probability and statistics model, and builds a series of feature templates according to the characteristics of objects themselves. Feature functions are generated based on the result of Chinese word segmentation and feature templates. It uses a limited memory BFGS algorithm to estimate parameters of the model, and labels property sequences of Web object blocks by Viterbi algorithm. Experiment result shows that the proposed method is an effective way to extract science data.

Original languageEnglish
Title of host publicationProceedings - 2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010
Pages483-487
Number of pages5
DOIs
StatePublished - 2010
Event2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010 - Chengdu, China
Duration: 9 Jul 201011 Jul 2010

Publication series

NameProceedings - 2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010
Volume4

Conference

Conference2010 3rd IEEE International Conference on Computer Science and Information Technology, ICCSIT 2010
Country/TerritoryChina
CityChengdu
Period9/07/1011/07/10

Keywords

  • Conditional random field
  • Information extraction
  • Machine learning
  • Web object

Fingerprint

Dive into the research topics of 'A CRF-based approach for web object extraction'. Together they form a unique fingerprint.

Cite this