跳到主要导航 跳到搜索 跳到主要内容

Text Similarity Measurement of Semantic Cognition Based on Word Vector Distance Decentralization with Clustering Analysis

  • Beihang University
  • North China University of Technology

科研成果: 期刊稿件文章同行评审

摘要

Text similarity measurement, which is a basic task in natural language processing, is widely used in text information mining, news classification and clustering, artificial intelligence, and other fields. This paper proposes a text similarity measure method named word vector distance decentralization (WVDD) which can deal with complex semantic relations, including sentence components, word order and weights for Chinese language. Then, the clustering analysis is performed for the obtained similarity results. A K-means algorithm based on Spark architecture for parallel computing is adopted to accelerate clustering speed here. In experimental verification, the test sets are significant number of customer comments posted on the Jingdong website, which is a comprehensive online shopping mall. F-measure is used to evaluate the accuracy of the results obtained by the proposed method. The superiority of the proposed method is verified and compared with the sentence vector model (Doc2vec) and bag-of-words model. The proposed method can be applied to analyze network language, such as customers' comments online and web chat data.

源语言英语
期刊论文编号8784295
页(从-至)107247-107258
页数12
期刊IEEE Access
7
DOI
出版状态已出版 - 2019

学术指纹

探究 'Text Similarity Measurement of Semantic Cognition Based on Word Vector Distance Decentralization with Clustering Analysis' 的科研主题。它们共同构成独一无二的学术指纹。

引用此