跳到主要导航 跳到搜索 跳到主要内容

大语言模型幻觉现象的分类识别与优化研究

  • Jing He
  • , Yang Shen*
  • , Runfeng Xie
  • *此作品的通讯作者
  • Tsinghua University
  • Beijing University of Technology

科研成果: 期刊稿件文章同行评审

摘要

With the widespread application of big language models in natural language understanding and generation tasks, their performance in high-precision fields such as healthcare, law, and scientific research has received increasing attention. However, the phenomenon of hallucinations, as a common problem in large language models, greatly restricts their practical application in these fields. At present, there are significant shortcomings in the evaluation and optimization of hallucination phenomena in large language models. Firstly, there is a lack of high-quality and high-precision domain hallucination evaluation datasets. Secondly, most of the existing hallucination assessment methods rely on a single model, which fails to take full advantage of the differences between multiple models. Finally, there are significant differences in the performance of different models in terms of hallucination types and rates, and there is currently no effective method to reduce the hallucination phenomenon in high hallucination rate models. This paper adopts a systematic process of dataset construction, swarm intelligence election, hallucination classification and quantification, and prior knowledge optimization to comprehensively evaluate and optimize the hallucination phenomenon of large language models in the field of medical question answering. Firstly, based on the publicly available dataset Huatuo, a large model illusion evaluation dataset in the medical question answering field is constructed by combining GPT generated question answers and manual annotation. Secondly, advanced big language models such as GPT4o, GPT4, ChatGLM4, Baichuan-13B, and Claude 3.5 are used to generate answers to questions in the dataset. By using a swarm intelligence based method, a LeaderAI is elected, which compares the answers of each model with reference answers to determine the illusion rate of each model. Finally, hallucinations are further divided into two categories: factual hallucinations and fidelity hallucinations. The research results indicate that under the guidance of LeaderAI, the illusion rate of the evaluated large models significantly decreases, especially the fidelity illusion rate.

投稿的翻译标题Research on Categorical Recognition and Optimization of Hallucination Phenomenon in Large Language Models
源语言繁体中文
页(从-至)1295-1301
页数7
期刊Journal of Frontiers of Computer Science and Technology
19
5
DOI
出版状态已出版 - 1 5月 2025

关键词

  • hallucination classification
  • hallucination recognition
  • large language model
  • model optimization

学术指纹

探究 '大语言模型幻觉现象的分类识别与优化研究' 的科研主题。它们共同构成独一无二的学术指纹。

引用此