CISum: Learning Cross-modality Interaction to Enhance Multimodal Semantic Coverage for Multimodal Summarization

  • Litian Zhang
  • , Xiaoming Zhang*
  • , Ziming Guo
  • , Zhipeng Liu
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Multimodal summarization (MS) aims to generate a summary from multimodal input. Previous works mainly focus on textual semantic coverage metrics such as ROUGE, which considers the visual content as supplemental data. Therefore, the summary is ineffective to cover the semantics of different modalities. This paper proposes a multi-task cross-modality learning framework (CISum) to improve multimodal semantic coverage by learning the cross-modality interaction in the multimodal article. To obtain the visual semantics, we translate images into visual descriptions based on the correlation with text content. Then, the visual description and text content are fused to generate the textual summary to capture the semantics of the multimodal content, and the most relevant image is selected as the visual summary. Furthermore, we design an automatic multimodal semantics coverage metric to evaluate the performance. Experimental results show that CISum outperforms baselines in multimodal semantics coverage metrics while maintaining the excellent performance of ROUGE and BLEU.

Original languageEnglish
Title of host publication2023 SIAM International Conference on Data Mining, SDM 2023
PublisherSociety for Industrial and Applied Mathematics Publications
Pages370-378
Number of pages9
ISBN (Electronic)9781611977653
DOIs
StatePublished - 2023
Event2023 SIAM International Conference on Data Mining, SDM 2023 - Minneapolis, United States
Duration: 27 Apr 202329 Apr 2023

Publication series

Name2023 SIAM International Conference on Data Mining, SDM 2023

Conference

Conference2023 SIAM International Conference on Data Mining, SDM 2023
Country/TerritoryUnited States
CityMinneapolis
Period27/04/2329/04/23

Keywords

  • Mulitmodal
  • Multi-task
  • Semantic coverage
  • Summarization

Fingerprint

Dive into the research topics of 'CISum: Learning Cross-modality Interaction to Enhance Multimodal Semantic Coverage for Multimodal Summarization'. Together they form a unique fingerprint.

Cite this