跳到主要导航 跳到搜索 跳到主要内容

HGTMFS: A Hypergraph Transformer Framework for Multimodal Summarization

  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Multimodal summarization, a rapidly evolving field within multimodal learning, focuses on generating cohesive summaries by integrating information from diverse modalities, such as text and images. Unlike traditional unimodal summarization, multimodal summarization presents unique challenges, particularly in capturing fine-grained interactions between modalities. Current models often fail to account for complex cross-modal interactions, leading to suboptimal performance and an over-reliance on one modality. To address these issues, we propose a novel framework, hypergraph transformer-based multimodal summarization (HGTMFS), designed to model high-order relationships across modalities. HGTMFS constructs a hypergraph that incorporates both textual and visual nodes and leverages transformer mechanisms to propagate information within the hypergraph. This approach enables the efficient exchange of multimodal data and improves the integration of fine-grained semantic relationships. Experimental results on several benchmark datasets demonstrate that HGTMFS outperforms state-of-the-art methods in multimodal summarization.

源语言英语
期刊论文编号9563
期刊Applied Sciences (Switzerland)
14
20
DOI
出版状态已出版 - 10月 2024

学术指纹

探究 'HGTMFS: A Hypergraph Transformer Framework for Multimodal Summarization' 的科研主题。它们共同构成独一无二的学术指纹。

引用此