Skip to main navigation Skip to search Skip to main content

CELB: Content extraction based on line-block

  • Xiao Ma*
  • , Jiangfeng Chen
  • , Hui Zhang
  • *Corresponding author for this work
  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper, we propose a simple, fast and accurate content extraction method: CELB. Compared with traditional methods, this approach does not parse the DOM trees and uses only information from lines of original HTML documents. We propose a concept called line-block, to extract contents more effectively and a new feature distance-text number (DTN) for distinctions between contents and non-contents. First, we preprocess original HTML documents, and then combine lines into line-blocks. Next, we calculate values of content features for each line-block, and use thresholds to determine whether a lineblock is part of the main content or not. Experiments show satisfied results, especially for the running time.

Original languageEnglish
Title of host publicationProceedings - 6th International Conference on Computer Sciences and Convergence Information Technology, ICCIT 2011
Pages412-417
Number of pages6
StatePublished - 2011
Event6th International Conference on Computer Sciences and Convergence Information Technology, ICCIT 2011 - Seogwipo, Jeju Island, Korea, Republic of
Duration: 29 Nov 20111 Dec 2011

Publication series

NameProceedings - 6th International Conference on Computer Sciences and Convergence Information Technology, ICCIT 2011

Conference

Conference6th International Conference on Computer Sciences and Convergence Information Technology, ICCIT 2011
Country/TerritoryKorea, Republic of
CitySeogwipo, Jeju Island
Period29/11/111/12/11

Fingerprint

Dive into the research topics of 'CELB: Content extraction based on line-block'. Together they form a unique fingerprint.

Cite this