TY - GEN
T1 - Research on Data Mining Methods in the Field of Quality Problem Analysis Based on BERT Model
AU - Ding, Zihuan
AU - Zhao, Guangyan
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024.
PY - 2024
Y1 - 2024
N2 - In the current quality problem analysis work, there are problems such as large amount of data with scattered distribution, isolated data and difficult machine understanding. Most of the current data mining work in the field is based on deep learning models, which is difficult to be integrated into the characteristics of the data in the field. Also, there are still some deficiencies in its accuracy rate and training speed. Therefore, this paper carries out the research of data mining methods in the field of quality problem analysis and incorporates the characteristics of data in the field on the basis of the existing model construction research. Focusing on the named entity recognition task and the relationship extraction task, the data are preprocessed by sequence annotation method to form the data sets of the two tasks. For the first task, a BERT-based recognition method is adopted, where the input is processed by word-level segmentation and the meaning features are learned. For the second task, a method also based on the BERT model is used. The models trained by the two tasks are used to achieve data mining work in the field of quality problem analysis. Comparative analysis by examples shows that the training results based on BERT model are better than those based on LSTM model in both tasks.
AB - In the current quality problem analysis work, there are problems such as large amount of data with scattered distribution, isolated data and difficult machine understanding. Most of the current data mining work in the field is based on deep learning models, which is difficult to be integrated into the characteristics of the data in the field. Also, there are still some deficiencies in its accuracy rate and training speed. Therefore, this paper carries out the research of data mining methods in the field of quality problem analysis and incorporates the characteristics of data in the field on the basis of the existing model construction research. Focusing on the named entity recognition task and the relationship extraction task, the data are preprocessed by sequence annotation method to form the data sets of the two tasks. For the first task, a BERT-based recognition method is adopted, where the input is processed by word-level segmentation and the meaning features are learned. For the second task, a method also based on the BERT model is used. The models trained by the two tasks are used to achieve data mining work in the field of quality problem analysis. Comparative analysis by examples shows that the training results based on BERT model are better than those based on LSTM model in both tasks.
KW - Data Mining
KW - Named Entity Recognition
KW - Quality Issues
KW - Relation Extraction
UR - https://www.scopus.com/pages/publications/85187683533
U2 - 10.1007/978-981-97-0837-6_8
DO - 10.1007/978-981-97-0837-6_8
M3 - 会议稿件
AN - SCOPUS:85187683533
SN - 9789819708369
T3 - Communications in Computer and Information Science
SP - 111
EP - 123
BT - Data Mining and Big Data - 8th International Conference, DMBD 2023, Proceedings
A2 - Tan, Ying
A2 - Shi, Yuhui
PB - Springer Science and Business Media Deutschland GmbH
T2 - 8th International Conference on Data Mining and Big Data, DMBD 2023
Y2 - 9 December 2023 through 12 December 2023
ER -