摘要
The accurate evaluation of the capabilities of a multiskilled dialogue system is important to satisfy the different demands of users, including social banter, profound knowledge-based discussions, role-playing conversations, and dialogue recommendations. Current benchmarks concentrate on assessing specific dialogue skills and cannot efficiently evaluate multiple dialogue skills concurrently. To facilitate the evaluation of multiskill dialogues, this study establishes a Chinese multiskill evaluation benchmark, which is the Multi-Skill Dialogue Evaluation Benchmark (MSDE). MSDE contains 1,781 dialogues and 21,218 utterances, which cover four common dialogue tasks: chit-chat, knowledge dialog, persona-based dialog, and dialog recommendations. We performed extensive experiments on MSDE and examined the correlation between automatic and human evaluation metrics. Results indicate that (1) among the four dialogue tasks, chit-chat is the most difficult to analyze, while knowledge dialogue is the easiest; (2) significant differences exist in the performance of various metrics on MSDE; (3) for human evaluation, the analysis complexity of each metric differs across varying dialogue tasks. Certain data will be made available on https://github.com/IRIP-LLM/MSDE, and all data will be released after sorting.
| 投稿的翻译标题 | Evaluation of Chinese multiskill dialogues |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 1281-1293 |
| 页数 | 13 |
| 期刊 | CAAI Transactions on Intelligent Systems |
| 卷 | 20 |
| 期 | 5 |
| DOI | |
| 出版状态 | 已出版 - 2025 |
关键词
- chit-chat; open domain dialogue
- conversational recommendation
- dialogue evaluation
- knowledge-grounded dialogue
- large language model
- multiskill dialogue
- persona-chat
学术指纹
探究 '中文多技能对话评估' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver