TY - GEN
T1 - Mis-prompt
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Zeng, Jiayi
AU - Feng, Yizhe
AU - He, Mengliang
AU - Lei, Wenhui
AU - Zhang, Wei
AU - Liu, Zeming
AU - Shi, Xiaoming
AU - Zhou, Aimin
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Large language models (LLMs) have demonstrated significant advancements in error handling. Current error-handling works are performed in a passive manner, with explicit error-handling instructions. However, in real-world scenarios, explicit error-handling instructions are usually unavailable. In this paper, our work identifies this challenge as how to conduct proactive error handling without explicit error handling instructions. To promote further research, this work introduces a new benchmark, termed Mis-prompt, consisting of four evaluation tasks, an error category taxonomy, and a new evaluation dataset. Furthermore, this work analyzes current LLMs' performance on the benchmark, and the experimental results reveal that current LLMs show poor performance on proactive error handling, and SFT on error handling instances improves LLMs' proactive error handling capabilities. Dataset and codes are available at https://github.com/Jiayi-Zeng/mis-prompt.
AB - Large language models (LLMs) have demonstrated significant advancements in error handling. Current error-handling works are performed in a passive manner, with explicit error-handling instructions. However, in real-world scenarios, explicit error-handling instructions are usually unavailable. In this paper, our work identifies this challenge as how to conduct proactive error handling without explicit error handling instructions. To promote further research, this work introduces a new benchmark, termed Mis-prompt, consisting of four evaluation tasks, an error category taxonomy, and a new evaluation dataset. Furthermore, this work analyzes current LLMs' performance on the benchmark, and the experimental results reveal that current LLMs show poor performance on proactive error handling, and SFT on error handling instances improves LLMs' proactive error handling capabilities. Dataset and codes are available at https://github.com/Jiayi-Zeng/mis-prompt.
UR - https://www.scopus.com/pages/publications/105021045729
U2 - 10.18653/v1/2025.acl-long.833
DO - 10.18653/v1/2025.acl-long.833
M3 - 会议稿件
AN - SCOPUS:105021045729
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 17007
EP - 17034
BT - Long Papers
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
Y2 - 27 July 2025 through 1 August 2025
ER -