TY - JOUR
T1 - HGTMFS
T2 - A Hypergraph Transformer Framework for Multimodal Summarization
AU - Lu, Ming
AU - Lu, Xinxi
AU - Zhang, Xiaoming
N1 - Publisher Copyright:
© 2024 by the authors.
PY - 2024/10
Y1 - 2024/10
N2 - Multimodal summarization, a rapidly evolving field within multimodal learning, focuses on generating cohesive summaries by integrating information from diverse modalities, such as text and images. Unlike traditional unimodal summarization, multimodal summarization presents unique challenges, particularly in capturing fine-grained interactions between modalities. Current models often fail to account for complex cross-modal interactions, leading to suboptimal performance and an over-reliance on one modality. To address these issues, we propose a novel framework, hypergraph transformer-based multimodal summarization (HGTMFS), designed to model high-order relationships across modalities. HGTMFS constructs a hypergraph that incorporates both textual and visual nodes and leverages transformer mechanisms to propagate information within the hypergraph. This approach enables the efficient exchange of multimodal data and improves the integration of fine-grained semantic relationships. Experimental results on several benchmark datasets demonstrate that HGTMFS outperforms state-of-the-art methods in multimodal summarization.
AB - Multimodal summarization, a rapidly evolving field within multimodal learning, focuses on generating cohesive summaries by integrating information from diverse modalities, such as text and images. Unlike traditional unimodal summarization, multimodal summarization presents unique challenges, particularly in capturing fine-grained interactions between modalities. Current models often fail to account for complex cross-modal interactions, leading to suboptimal performance and an over-reliance on one modality. To address these issues, we propose a novel framework, hypergraph transformer-based multimodal summarization (HGTMFS), designed to model high-order relationships across modalities. HGTMFS constructs a hypergraph that incorporates both textual and visual nodes and leverages transformer mechanisms to propagate information within the hypergraph. This approach enables the efficient exchange of multimodal data and improves the integration of fine-grained semantic relationships. Experimental results on several benchmark datasets demonstrate that HGTMFS outperforms state-of-the-art methods in multimodal summarization.
KW - hypergraph transformer
KW - multimodal summarization
UR - https://www.scopus.com/pages/publications/85207477633
U2 - 10.3390/app14209563
DO - 10.3390/app14209563
M3 - 文章
AN - SCOPUS:85207477633
SN - 2076-3417
VL - 14
JO - Applied Sciences (Switzerland)
JF - Applied Sciences (Switzerland)
IS - 20
M1 - 9563
ER -