跳到主要导航 跳到搜索 跳到主要内容

ANEEC: A quasi-automatic system for massive named entity extraction and categorization

  • Bingyue Peng
  • , Junjie Wu
  • , Hua Yuan
  • , Qingwei Guo
  • , Dacheng Tao*
  • *此作品的通讯作者
  • Beihang University
  • University of Electronic Science and Technology of China
  • Microsoft USA
  • University of Technology Sydney

科研成果: 期刊稿件文章同行评审

摘要

Named entity recognition seeks to locate atomic elements in texts and classify them into predefined categories. It is essentially useful for many applications, including microblog analysis and query suggestion. In recent years, with the explosion of Web 2.0, people have found it a promising way to extract large-scale, high-quality entities from structured web content. However, existing studies seldom provide an integrated system for simultaneously extracting and categorizing both the head and tail entities, and the identification of ambiguous entities is still a challenging task. In light of these, we propose a system named quasi-Automatic Named Entity Extraction and Categorization (ANEEC) for massive named-entity management. Specifically, ANEECfirst identifies representative websites by using a small seed-set of entities and the query logs of a search engine, and then extracts high-quality entities from the parallel structures in the webpages. ANEEC then employs the extracted entities and their corresponding atom-level groups to establish an entity taxonomy as well as a hierarchical classifier ensemble. Two problems, i.e. definition abnormality and granularity unfitness, have also been addressed to further improve the quality of the taxonomy. An application case using 932 seed entities and the query logs of the search engine Bing demonstrates that ANEEC can effectively identify over 870 000 named entities in 32 bottom-level categories, and the resulting taxonomy has an excellent classification performance with F1 = 85.17%, provided that the entity features are properly preprocessed and weighted. In particular,ANEEC shows the potential for tail entity recognition and ambiguous entity detection.

源语言英语
页(从-至)1328-1346
页数19
期刊Computer Journal
56
11
DOI
出版状态已出版 - 11月 2013

学术指纹

探究 'ANEEC: A quasi-automatic system for massive named entity extraction and categorization' 的科研主题。它们共同构成独一无二的学术指纹。

引用此