跳到主要导航 跳到搜索 跳到主要内容

Revisiting audio visual scene-aware dialog

  • University of Cambridge
  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Audio Visual Scene-Aware Dialog (AVSD) has drawn intense interests, in which models are required to understand dynamic scenes in videos and dialog contexts in order to converse with human users by generating responses to given questions. Existing works have laid a solid foundation towards solving the AVSD problem. In contrast to previous studies, this paper empirically revisits the AVSD task and argues that this task exhibits a variety of biases in terms of models, dataset, and evaluation metrics: (1) as for the models, we believe that the state-of-the-art frameworks do not utilize multimodal features to their full extent; (2) as for the dataset, we conduct a deep analysis into dataset statistics from different types of questions and find that the dataset is slightly biased in several specific aspects; by simply implementing a caption-only baseline that has never seen the video, we achieve state-of-the-art performance on the AVSD task; (3) as for the evaluation metrics, we argue that the current metrics for AVSD primarily focus on the naturalness of generated responses while ignoring the truthfulness, which makes them fall short of disclosing the consistency of model predictions and the actual visual content. Overall, our analysis aims to provide a detailed inspection of the AVSD task and we hope that our empirical observations can inspire further improvement to the task.

源语言英语
页(从-至)227-237
页数11
期刊Neurocomputing
496
DOI
出版状态已出版 - 28 7月 2022

学术指纹

探究 'Revisiting audio visual scene-aware dialog' 的科研主题。它们共同构成独一无二的学术指纹。

引用此