跳到主要导航 跳到搜索 跳到主要内容

Toward Generalized Multimodal Sarcasm Detection With Multi-View Learning

  • Diandian Guo
  • , Hao Peng
  • , Cong Cao*
  • , Fangfang Yuan
  • , Yanbing Liu*
  • , Philip S. Yu
  • *此作品的通讯作者
  • CAS - Institute of Information Engineering
  • University of Chinese Academy of Sciences
  • University of Illinois at Chicago

科研成果: 期刊稿件文章同行评审

摘要

Multimodal Sarcasm Detection (MSD) is crucial for understanding complex human communication and building intelligent emotional computing systems. However, existing MSD methods often over-rely on spurious correlations, causing the learned features to deviate from the true semantics of sarcasm. This bias significantly undermines the generalization ability of current models beyond the training environment. This paper proposes a Confict-based Disentangled multi-view incongruity learning framework (ConDi), aiming to effectively disentangle and interact with heterogeneous features in multimodal sarcasm. Given that multimodal embedding spaces are typically heterogeneous, direct fusion can disrupt the inherent structure of embeddings from different modalities. To address this issue, we employ the optimal transport algorithm to align embeddings from different modalities into a unified space. Subsequently, we jointly learn incongruities from three views: modality disentanglement, global sentiment, and local description. To achieve a debiased fusion of sarcasm features, we design a conflict-based fusion module to integrate features from these three views. Experimental results demonstrate the superiority of ConDi on multimodal sarcasm datasets, and further analysis shows that ConDi can effectively reduce reliance on spurious correlations. Additionally, out-of-distribution (OOD) experiments reveal that ConDi achieves better generalization.

源语言英语
页(从-至)4995-5009
页数15
期刊IEEE Transactions on Audio, Speech and Language Processing
33
DOI
出版状态已出版 - 2025

指纹

探究 'Toward Generalized Multimodal Sarcasm Detection With Multi-View Learning' 的科研主题。它们共同构成独一无二的指纹。

引用此