TY - JOUR
T1 - Revisiting audio visual scene-aware dialog
AU - Liu, Aishan
AU - Xie, Huiyuan
AU - Liu, Xianglong
AU - Yin, Zixin
AU - Liu, Shunchang
N1 - Publisher Copyright:
© 2022 Elsevier B.V.
PY - 2022/7/28
Y1 - 2022/7/28
N2 - Audio Visual Scene-Aware Dialog (AVSD) has drawn intense interests, in which models are required to understand dynamic scenes in videos and dialog contexts in order to converse with human users by generating responses to given questions. Existing works have laid a solid foundation towards solving the AVSD problem. In contrast to previous studies, this paper empirically revisits the AVSD task and argues that this task exhibits a variety of biases in terms of models, dataset, and evaluation metrics: (1) as for the models, we believe that the state-of-the-art frameworks do not utilize multimodal features to their full extent; (2) as for the dataset, we conduct a deep analysis into dataset statistics from different types of questions and find that the dataset is slightly biased in several specific aspects; by simply implementing a caption-only baseline that has never seen the video, we achieve state-of-the-art performance on the AVSD task; (3) as for the evaluation metrics, we argue that the current metrics for AVSD primarily focus on the naturalness of generated responses while ignoring the truthfulness, which makes them fall short of disclosing the consistency of model predictions and the actual visual content. Overall, our analysis aims to provide a detailed inspection of the AVSD task and we hope that our empirical observations can inspire further improvement to the task.
AB - Audio Visual Scene-Aware Dialog (AVSD) has drawn intense interests, in which models are required to understand dynamic scenes in videos and dialog contexts in order to converse with human users by generating responses to given questions. Existing works have laid a solid foundation towards solving the AVSD problem. In contrast to previous studies, this paper empirically revisits the AVSD task and argues that this task exhibits a variety of biases in terms of models, dataset, and evaluation metrics: (1) as for the models, we believe that the state-of-the-art frameworks do not utilize multimodal features to their full extent; (2) as for the dataset, we conduct a deep analysis into dataset statistics from different types of questions and find that the dataset is slightly biased in several specific aspects; by simply implementing a caption-only baseline that has never seen the video, we achieve state-of-the-art performance on the AVSD task; (3) as for the evaluation metrics, we argue that the current metrics for AVSD primarily focus on the naturalness of generated responses while ignoring the truthfulness, which makes them fall short of disclosing the consistency of model predictions and the actual visual content. Overall, our analysis aims to provide a detailed inspection of the AVSD task and we hope that our empirical observations can inspire further improvement to the task.
KW - Modality bias
KW - Multimodal dialog systems
KW - Multimodal evaluation
UR - https://www.scopus.com/pages/publications/85124745887
U2 - 10.1016/j.neucom.2021.08.151
DO - 10.1016/j.neucom.2021.08.151
M3 - 文章
AN - SCOPUS:85124745887
SN - 0925-2312
VL - 496
SP - 227
EP - 237
JO - Neurocomputing
JF - Neurocomputing
ER -