跳到主要导航 跳到搜索 跳到主要内容

Learning fine-grained representation with token-level alignment for multimodal sentiment analysis

  • Xiang Li
  • , Haijun Zhang
  • , Zhiqiang Dong
  • , Xianfu Cheng
  • , Yun Liu
  • , Xiaoming Zhang*
  • *此作品的通讯作者
  • Beihang University
  • Beijing Wuzi University
  • CSSC Systems Engineering Research Institute
  • Moutai Institute

科研成果: 期刊稿件文章同行评审

摘要

Multimodal sentiment analysis (MSA) has gained significant attention recently due to its crucial applications in intelligent systems. The effective extraction and fusion of sentiment representations from heterogeneous modalities (visual, text, and audio) remain the key and challenge in improving MSA performance. However, most existing methods often directly integrate different modalities and fail to capture fine-grained multimodal representations, suffering from irrelevant information among heterogeneous modalities. Moreover, these methods overlook the explicit token-level alignment between different modalities, which hinders effective representation learning and multimodal fusion. To address these issues, we propose a Fine-grained Multimodal Fusion Network (FMFN) for MSA. Firstly, a Fine-grained Representation Learning (FRL) module is designed to extract critical sentiment information from original modality features by employing finite learnable denoising tokens. FRL enables sentiment-related fine-grained representations to interact with these denoising tokens, thereby mitigating noise propagation and information redundancy. Secondly, a Token-level Cross-modal Alignment (TCA) module is introduced, which aligns different modality representations through fine-grained contrastive learning. TCA refines each modality by filtering out extraneous representations while facilitating subsequent multimodal interactions. Finally, a Correlation-aware Multimodal Fusion (CMF) module is developed to capture latent cross-modal correlations, generating consistent and complementary multimodal representations that enhance sentiment prediction. Extensive experiments are conducted on two English datasets, CMU-MOSI and CMU-MOSEI (aligned and unaligned versions), and one Chinese dataset, CH-SIMS (unaligned). The proposed method outperforms state-of-the-art models with 1%–2% accuracy and F1 score improvements, particularly on unaligned datasets, demonstrating the effectiveness of learning the fine-grained representation and token-level alignment in enhancing MSA performance.

源语言英语
文章编号126274
期刊Expert Systems with Applications
269
DOI
出版状态已出版 - 15 4月 2025

指纹

探究 'Learning fine-grained representation with token-level alignment for multimodal sentiment analysis' 的科研主题。它们共同构成独一无二的指纹。

引用此