Skip to main navigation Skip to search Skip to main content

Learning fine-grained representation with token-level alignment for multimodal sentiment analysis

  • Xiang Li
  • , Haijun Zhang
  • , Zhiqiang Dong
  • , Xianfu Cheng
  • , Yun Liu
  • , Xiaoming Zhang*
  • *Corresponding author for this work
  • Beihang University
  • Beijing Wuzi University
  • CSSC Systems Engineering Research Institute
  • Moutai Institute

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal sentiment analysis (MSA) has gained significant attention recently due to its crucial applications in intelligent systems. The effective extraction and fusion of sentiment representations from heterogeneous modalities (visual, text, and audio) remain the key and challenge in improving MSA performance. However, most existing methods often directly integrate different modalities and fail to capture fine-grained multimodal representations, suffering from irrelevant information among heterogeneous modalities. Moreover, these methods overlook the explicit token-level alignment between different modalities, which hinders effective representation learning and multimodal fusion. To address these issues, we propose a Fine-grained Multimodal Fusion Network (FMFN) for MSA. Firstly, a Fine-grained Representation Learning (FRL) module is designed to extract critical sentiment information from original modality features by employing finite learnable denoising tokens. FRL enables sentiment-related fine-grained representations to interact with these denoising tokens, thereby mitigating noise propagation and information redundancy. Secondly, a Token-level Cross-modal Alignment (TCA) module is introduced, which aligns different modality representations through fine-grained contrastive learning. TCA refines each modality by filtering out extraneous representations while facilitating subsequent multimodal interactions. Finally, a Correlation-aware Multimodal Fusion (CMF) module is developed to capture latent cross-modal correlations, generating consistent and complementary multimodal representations that enhance sentiment prediction. Extensive experiments are conducted on two English datasets, CMU-MOSI and CMU-MOSEI (aligned and unaligned versions), and one Chinese dataset, CH-SIMS (unaligned). The proposed method outperforms state-of-the-art models with 1%–2% accuracy and F1 score improvements, particularly on unaligned datasets, demonstrating the effectiveness of learning the fine-grained representation and token-level alignment in enhancing MSA performance.

Original languageEnglish
Article number126274
JournalExpert Systems with Applications
Volume269
DOIs
StatePublished - 15 Apr 2025

Keywords

  • Fine-grained representation
  • Multimodal fusion
  • Representation learning
  • Sentiment analysis
  • Token-level alignment

Fingerprint

Dive into the research topics of 'Learning fine-grained representation with token-level alignment for multimodal sentiment analysis'. Together they form a unique fingerprint.

Cite this