Skip to main navigation Skip to search Skip to main content

Toward Generalized Multimodal Sarcasm Detection With Multi-View Learning

  • Diandian Guo
  • , Hao Peng
  • , Cong Cao*
  • , Fangfang Yuan
  • , Yanbing Liu*
  • , Philip S. Yu
  • *Corresponding author for this work
  • CAS - Institute of Information Engineering
  • University of Chinese Academy of Sciences
  • University of Illinois at Chicago

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal Sarcasm Detection (MSD) is crucial for understanding complex human communication and building intelligent emotional computing systems. However, existing MSD methods often over-rely on spurious correlations, causing the learned features to deviate from the true semantics of sarcasm. This bias significantly undermines the generalization ability of current models beyond the training environment. This paper proposes a Confict-based Disentangled multi-view incongruity learning framework (ConDi), aiming to effectively disentangle and interact with heterogeneous features in multimodal sarcasm. Given that multimodal embedding spaces are typically heterogeneous, direct fusion can disrupt the inherent structure of embeddings from different modalities. To address this issue, we employ the optimal transport algorithm to align embeddings from different modalities into a unified space. Subsequently, we jointly learn incongruities from three views: modality disentanglement, global sentiment, and local description. To achieve a debiased fusion of sarcasm features, we design a conflict-based fusion module to integrate features from these three views. Experimental results demonstrate the superiority of ConDi on multimodal sarcasm datasets, and further analysis shows that ConDi can effectively reduce reliance on spurious correlations. Additionally, out-of-distribution (OOD) experiments reveal that ConDi achieves better generalization.

Original languageEnglish
Pages (from-to)4995-5009
Number of pages15
JournalIEEE Transactions on Audio, Speech and Language Processing
Volume33
DOIs
StatePublished - 2025

Keywords

  • debiased learning
  • generalization
  • multimodal
  • sarcasm detection
  • spurious correlation

Fingerprint

Dive into the research topics of 'Toward Generalized Multimodal Sarcasm Detection With Multi-View Learning'. Together they form a unique fingerprint.

Cite this