Skip to main navigation Skip to search Skip to main content

Learning joint multimodal representation with adversarial attention networks

  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recently, learning a joint representation for the multimodal data (e.g., containing both visual content and text description) has attracted extensive research interests. Usually, the features of different modalities are correlational and compositive, and thus a joint representation capturing the correlation is more effective than a subset of the features. Most of existing multimodal representation learning methods suffer from lack of additional constraints to enhance the robustness of the learned representations. In this paper, a novel Adversarial Attention Networks (AAN) is proposed to incorporate both the attention mechanism and the adversarial networks for effective and robust multimodal representation learning. Specifically, a visual-semantic attention model with siamese learning strategy is proposed to encode the fine-grained correlation between visual and textual modalities. Meanwhile, the adversarial learning model is employed to regularize the generated representation by matching the posterior distribution of the representation to the given priors. Then, the two modules are incorporated into a integrated learning framework to learn the joint multimodal representation. Experimental results in two tasks, i.e., multi-label classification and tag recommendation, show that the proposed model outperforms state-of-the-art representation learning methods.

Original languageEnglish
Title of host publicationMM 2018 - Proceedings of the 2018 ACM Multimedia Conference
PublisherAssociation for Computing Machinery, Inc
Pages1874-1882
Number of pages9
ISBN (Electronic)9781450356657
DOIs
StatePublished - 15 Oct 2018
Event26th ACM Multimedia conference, MM 2018 - Seoul, Korea, Republic of
Duration: 22 Oct 201826 Oct 2018

Publication series

NameMM 2018 - Proceedings of the 2018 ACM Multimedia Conference

Conference

Conference26th ACM Multimedia conference, MM 2018
Country/TerritoryKorea, Republic of
CitySeoul
Period22/10/1826/10/18

Keywords

  • Adversarial networks
  • Attention model
  • Multimodal
  • Representation learning
  • Siamese learning

Fingerprint

Dive into the research topics of 'Learning joint multimodal representation with adversarial attention networks'. Together they form a unique fingerprint.

Cite this