跳到主要导航 跳到搜索 跳到主要内容

Multimodal feature fusion by relational reasoning and attention for visual question answering

  • Weifeng Zhang*
  • , Jing Yu
  • , Hua Hu
  • , Haiyang Hu
  • , Zengchang Qin
  • *此作品的通讯作者
  • Jiaxing University
  • University of Chinese Academy of Sciences
  • Hangzhou Dianzi University

科研成果: 期刊稿件文章同行评审

摘要

The recently emerged research of Visual Question Answering (VQA) has become a hot topic in computer vision. A key solution to VQA exists in how to fuse multimodal features extracted from image and question. In this paper, we show that combining visual relationship and attention together achieves more fine-grained feature fusion. Specifically, we design an effective and efficient module to reason complex relationship between visual objects. In addition, a bilinear attention module is learned for question guided attention on visual objects, which allows us to obtain more discriminative visual features. Given an image and a question in natural language, our VQA model learns visual relational reasoning network and attention network in parallel to fuse fine-grained textual and visual features, so that answers can be predicted accurately. Experimental results show that our approach achieves new state-of-the-art performance of single model on both VQA 1.0 and VQA 2.0 datasets.

源语言英语
页(从-至)116-126
页数11
期刊Information Fusion
55
DOI
出版状态已出版 - 3月 2020

指纹

探究 'Multimodal feature fusion by relational reasoning and attention for visual question answering' 的科研主题。它们共同构成独一无二的指纹。

引用此