跳到主要导航 跳到搜索 跳到主要内容

Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling

  • Jiayi Zeng
  • , Yizhe Feng
  • , Mengliang He
  • , Wenhui Lei
  • , Wei Zhang
  • , Zeming Liu*
  • , Xiaoming Shi*
  • , Aimin Zhou
  • *此作品的通讯作者
  • East China Normal University
  • Beihang University
  • Shanghai Jiao Tong University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Large language models (LLMs) have demonstrated significant advancements in error handling. Current error-handling works are performed in a passive manner, with explicit error-handling instructions. However, in real-world scenarios, explicit error-handling instructions are usually unavailable. In this paper, our work identifies this challenge as how to conduct proactive error handling without explicit error handling instructions. To promote further research, this work introduces a new benchmark, termed Mis-prompt, consisting of four evaluation tasks, an error category taxonomy, and a new evaluation dataset. Furthermore, this work analyzes current LLMs' performance on the benchmark, and the experimental results reveal that current LLMs show poor performance on proactive error handling, and SFT on error handling instances improves LLMs' proactive error handling capabilities. Dataset and codes are available at https://github.com/Jiayi-Zeng/mis-prompt.

源语言英语
主期刊名Long Papers
编辑Wanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
出版商Association for Computational Linguistics (ACL)
17007-17034
页数28
ISBN(电子版)9798891762510
DOI
出版状态已出版 - 2025
活动63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025 - Vienna, 奥地利
期限: 27 7月 20251 8月 2025

出版系列

姓名Proceedings of the Annual Meeting of the Association for Computational Linguistics
1
ISSN(印刷版)0736-587X

会议

会议63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
国家/地区奥地利
Vienna
时期27/07/251/08/25

指纹

探究 'Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling' 的科研主题。它们共同构成独一无二的指纹。

引用此