跳到主要导航 跳到搜索 跳到主要内容

Vision–language representation learning with breadth and depth attention pre-training

  • Yun Liu*
  • , Bo Zhang
  • , Chen Cheng Wang
  • , Genglong Yan
  • , Ke Zhou
  • , Zhoujun Li
  • , Lei Lei Zhang
  • *此作品的通讯作者
  • Moutai Institute
  • Nanjing Normal University

科研成果: 期刊稿件文章同行评审

摘要

The rapid advances in computer vision and natural language processing have led to increased attention toward the challenge of understanding vision and language together across multiple domains. Representation learning has become a major focus of research on cross-modal information understanding. However, current methods often fall short of providing comprehensive interaction and meaningful supervised guidance that would allow for effective learning of visual-linguistic joint representation. In this paper, we introduce the Breadth and Depth Attention Pre-training (BDAP) model for vision–language representation learning. Our model includes a breadth attention network designed to model feature associations between text sentences and image regions across different image levels. It uses fine-grained image features to promote more effective cross-modal feature interactions. Additionally, a depth attention network, which repeatedly calculates attention scores, is designed to deeply capture the complementarity between the image and text by gradually refining important image regions related to the text. Furthermore, we propose an attention pre-training network that leverages attention annotated distribution maps as prior knowledge to supervise the learning process of the breadth and depth attention networks, thereby enabling weight initialization of both types of attention networks. Extensive experiments on datasets of visual question answering and multi-modal sentiment analysis demonstrate the promising superiority of our BDAP model for vision–language representation learning.

源语言英语
期刊论文编号112941
期刊Knowledge-Based Systems
310
DOI
出版状态已出版 - 15 2月 2025

学术指纹

探究 'Vision–language representation learning with breadth and depth attention pre-training' 的科研主题。它们共同构成独一无二的学术指纹。

引用此