Skip to main navigation Skip to search Skip to main content

Vision–language representation learning with breadth and depth attention pre-training

  • Yun Liu*
  • , Bo Zhang
  • , Chen Cheng Wang
  • , Genglong Yan
  • , Ke Zhou
  • , Zhoujun Li
  • , Lei Lei Zhang
  • *Corresponding author for this work
  • Moutai Institute
  • Nanjing Normal University

Research output: Contribution to journalArticlepeer-review

Abstract

The rapid advances in computer vision and natural language processing have led to increased attention toward the challenge of understanding vision and language together across multiple domains. Representation learning has become a major focus of research on cross-modal information understanding. However, current methods often fall short of providing comprehensive interaction and meaningful supervised guidance that would allow for effective learning of visual-linguistic joint representation. In this paper, we introduce the Breadth and Depth Attention Pre-training (BDAP) model for vision–language representation learning. Our model includes a breadth attention network designed to model feature associations between text sentences and image regions across different image levels. It uses fine-grained image features to promote more effective cross-modal feature interactions. Additionally, a depth attention network, which repeatedly calculates attention scores, is designed to deeply capture the complementarity between the image and text by gradually refining important image regions related to the text. Furthermore, we propose an attention pre-training network that leverages attention annotated distribution maps as prior knowledge to supervise the learning process of the breadth and depth attention networks, thereby enabling weight initialization of both types of attention networks. Extensive experiments on datasets of visual question answering and multi-modal sentiment analysis demonstrate the promising superiority of our BDAP model for vision–language representation learning.

Original languageEnglish
Article number112941
JournalKnowledge-Based Systems
Volume310
DOIs
StatePublished - 15 Feb 2025

Keywords

  • Attention
  • Pre-training
  • Representation learning

Fingerprint

Dive into the research topics of 'Vision–language representation learning with breadth and depth attention pre-training'. Together they form a unique fingerprint.

Cite this