TY - JOUR
T1 - Zero-shot reinforcement learning for multi-domain task-oriented dialogue policy
AU - Ren, Yuan
AU - Chen, Si
AU - Zhang, Ri Chong
AU - Liu, Xu Dong
AU - Peng, Ming Tian
N1 - Publisher Copyright:
© Higher Education Press 2026.
PY - 2026/9
Y1 - 2026/9
N2 - Dialogue policy, a critical component in multi-domain task-oriented dialogue systems, decides the dialogue acts according to the received dialogue state. We introduce zero-shot reinforcement learning for dialogue policy learning, which aims to learn dialogue policies capable of generalizing to unseen domains without further training. This setup brings forward two challenges: 1) the representation of unseen actions & states, and 2) zero-shot generalization to unseen domains. For the first issue, we propose Unified Representation (UR), an ontology-agnostic representation, which effectively infers representations in unseen domains by capturing the underlying semantic relations between unseen actions and states and seen ones. To tackle the second issue, we propose Q-Values Perturbation (QVP), a family of exploration strategies that can be applied either during training or testing. Experiments on MultiWOZ, suggest that UR, QVP, and an integrated framework combining the two are all effective.
AB - Dialogue policy, a critical component in multi-domain task-oriented dialogue systems, decides the dialogue acts according to the received dialogue state. We introduce zero-shot reinforcement learning for dialogue policy learning, which aims to learn dialogue policies capable of generalizing to unseen domains without further training. This setup brings forward two challenges: 1) the representation of unseen actions & states, and 2) zero-shot generalization to unseen domains. For the first issue, we propose Unified Representation (UR), an ontology-agnostic representation, which effectively infers representations in unseen domains by capturing the underlying semantic relations between unseen actions and states and seen ones. To tackle the second issue, we propose Q-Values Perturbation (QVP), a family of exploration strategies that can be applied either during training or testing. Experiments on MultiWOZ, suggest that UR, QVP, and an integrated framework combining the two are all effective.
KW - dialogue system
KW - generalization
KW - reinforcement learning
KW - unified representation
KW - zero-Shot
UR - https://www.scopus.com/pages/publications/105029838384
U2 - 10.1007/s11704-025-41285-5
DO - 10.1007/s11704-025-41285-5
M3 - 文章
AN - SCOPUS:105029838384
SN - 2095-2228
VL - 20
JO - Frontiers of Computer Science
JF - Frontiers of Computer Science
IS - 9
M1 - 2009353
ER -