Skip to main navigation Skip to search Skip to main content

Toward multimodal sentiment analysis with a self-supervised knowledge-augmented network

  • Yun Liu
  • , Xiaoming Zhang
  • , Tianhao Peng*
  • , Ke Zhou
  • , Zhoujun Li
  • *Corresponding author for this work
  • Moutai Institute
  • Beihang University

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal sentiment analysis (MSA) has attracted increasing attention for its ability to exploit complementary emotional cues from multiple modalities. However, existing methods still encounter two critical limitations:(1) Overemphasis on cross-modal alignment while neglecting in-depth analysis of emotion-specific cross-modal interaction cues, and (2) Reliance on limited labeled data, leading to overfitting in supervised models. To address these challenges, this paper proposes SKAN: A Self-supervised Knowledge-Augmented Network for Multimodal Sentiment Analysis. First, multimodal information is input to a large vision-language model to generate explicit cross-modal sentiment descriptions. The sentiment descriptions, acting as external knowledge, are integrated with the corresponding text-image pairs through a text-centric multimodal fusion module. It augments the model’s ability to discover latent sentiment correlations and improves multimodal sentiment expression capabilities. Second, to alleviate the impact of data scarcity, a self-supervised pretraining strategy is devised, leveraging a sentiment intensity lexicon to perform emotion masking and intensity estimation on unlabeled multimodal data. This design enables the model to acquire cross-modal emotional representations from vast unlabeled samples, thereby improving its semantic sensitivity and generalization ability. Extensive experiments on three benchmark datasets validate the superior performance of SKAN compared with state-of-the-art baselines. The proposed framework provides a novel paradigm that synergistically integrates external knowledge and self-supervision to advance the field of multimodal sentiment analysis.

Original languageEnglish
Article number131648
JournalExpert Systems with Applications
Volume314
DOIs
StatePublished - 5 Jun 2026

Keywords

  • Knowledge-augmented network
  • Multimodal sentiment analysis
  • Self-supervised learning

Fingerprint

Dive into the research topics of 'Toward multimodal sentiment analysis with a self-supervised knowledge-augmented network'. Together they form a unique fingerprint.

Cite this