TY - JOUR
T1 - ANEEC
T2 - A quasi-automatic system for massive named entity extraction and categorization
AU - Peng, Bingyue
AU - Wu, Junjie
AU - Yuan, Hua
AU - Guo, Qingwei
AU - Tao, Dacheng
PY - 2013/11
Y1 - 2013/11
N2 - Named entity recognition seeks to locate atomic elements in texts and classify them into predefined categories. It is essentially useful for many applications, including microblog analysis and query suggestion. In recent years, with the explosion of Web 2.0, people have found it a promising way to extract large-scale, high-quality entities from structured web content. However, existing studies seldom provide an integrated system for simultaneously extracting and categorizing both the head and tail entities, and the identification of ambiguous entities is still a challenging task. In light of these, we propose a system named quasi-Automatic Named Entity Extraction and Categorization (ANEEC) for massive named-entity management. Specifically, ANEECfirst identifies representative websites by using a small seed-set of entities and the query logs of a search engine, and then extracts high-quality entities from the parallel structures in the webpages. ANEEC then employs the extracted entities and their corresponding atom-level groups to establish an entity taxonomy as well as a hierarchical classifier ensemble. Two problems, i.e. definition abnormality and granularity unfitness, have also been addressed to further improve the quality of the taxonomy. An application case using 932 seed entities and the query logs of the search engine Bing demonstrates that ANEEC can effectively identify over 870 000 named entities in 32 bottom-level categories, and the resulting taxonomy has an excellent classification performance with F1 = 85.17%, provided that the entity features are properly preprocessed and weighted. In particular,ANEEC shows the potential for tail entity recognition and ambiguous entity detection.
AB - Named entity recognition seeks to locate atomic elements in texts and classify them into predefined categories. It is essentially useful for many applications, including microblog analysis and query suggestion. In recent years, with the explosion of Web 2.0, people have found it a promising way to extract large-scale, high-quality entities from structured web content. However, existing studies seldom provide an integrated system for simultaneously extracting and categorizing both the head and tail entities, and the identification of ambiguous entities is still a challenging task. In light of these, we propose a system named quasi-Automatic Named Entity Extraction and Categorization (ANEEC) for massive named-entity management. Specifically, ANEECfirst identifies representative websites by using a small seed-set of entities and the query logs of a search engine, and then extracts high-quality entities from the parallel structures in the webpages. ANEEC then employs the extracted entities and their corresponding atom-level groups to establish an entity taxonomy as well as a hierarchical classifier ensemble. Two problems, i.e. definition abnormality and granularity unfitness, have also been addressed to further improve the quality of the taxonomy. An application case using 932 seed entities and the query logs of the search engine Bing demonstrates that ANEEC can effectively identify over 870 000 named entities in 32 bottom-level categories, and the resulting taxonomy has an excellent classification performance with F1 = 85.17%, provided that the entity features are properly preprocessed and weighted. In particular,ANEEC shows the potential for tail entity recognition and ambiguous entity detection.
KW - Entity categorization
KW - Entity disambiguation
KW - Entity extraction
KW - Named entity
KW - Structured web
UR - https://www.scopus.com/pages/publications/84889845373
U2 - 10.1093/comjnl/bxs114
DO - 10.1093/comjnl/bxs114
M3 - 文章
AN - SCOPUS:84889845373
SN - 0010-4620
VL - 56
SP - 1328
EP - 1346
JO - Computer Journal
JF - Computer Journal
IS - 11
ER -