跳到主要导航 跳到搜索 跳到主要内容

CISum: Learning Cross-modality Interaction to Enhance Multimodal Semantic Coverage for Multimodal Summarization

  • Litian Zhang
  • , Xiaoming Zhang*
  • , Ziming Guo
  • , Zhipeng Liu
  • *此作品的通讯作者
  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Multimodal summarization (MS) aims to generate a summary from multimodal input. Previous works mainly focus on textual semantic coverage metrics such as ROUGE, which considers the visual content as supplemental data. Therefore, the summary is ineffective to cover the semantics of different modalities. This paper proposes a multi-task cross-modality learning framework (CISum) to improve multimodal semantic coverage by learning the cross-modality interaction in the multimodal article. To obtain the visual semantics, we translate images into visual descriptions based on the correlation with text content. Then, the visual description and text content are fused to generate the textual summary to capture the semantics of the multimodal content, and the most relevant image is selected as the visual summary. Furthermore, we design an automatic multimodal semantics coverage metric to evaluate the performance. Experimental results show that CISum outperforms baselines in multimodal semantics coverage metrics while maintaining the excellent performance of ROUGE and BLEU.

源语言英语
主期刊名2023 SIAM International Conference on Data Mining, SDM 2023
出版商Society for Industrial and Applied Mathematics Publications
370-378
页数9
ISBN(电子版)9781611977653
DOI
出版状态已出版 - 2023
活动2023 SIAM International Conference on Data Mining, SDM 2023 - Minneapolis, 美国
期限: 27 4月 202329 4月 2023

出版系列

姓名2023 SIAM International Conference on Data Mining, SDM 2023

会议

会议2023 SIAM International Conference on Data Mining, SDM 2023
国家/地区美国
Minneapolis
时期27/04/2329/04/23

指纹

探究 'CISum: Learning Cross-modality Interaction to Enhance Multimodal Semantic Coverage for Multimodal Summarization' 的科研主题。它们共同构成独一无二的指纹。

引用此