TY - JOUR
T1 - Learning fine-grained representation with token-level alignment for multimodal sentiment analysis
AU - Li, Xiang
AU - Zhang, Haijun
AU - Dong, Zhiqiang
AU - Cheng, Xianfu
AU - Liu, Yun
AU - Zhang, Xiaoming
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2025/4/15
Y1 - 2025/4/15
N2 - Multimodal sentiment analysis (MSA) has gained significant attention recently due to its crucial applications in intelligent systems. The effective extraction and fusion of sentiment representations from heterogeneous modalities (visual, text, and audio) remain the key and challenge in improving MSA performance. However, most existing methods often directly integrate different modalities and fail to capture fine-grained multimodal representations, suffering from irrelevant information among heterogeneous modalities. Moreover, these methods overlook the explicit token-level alignment between different modalities, which hinders effective representation learning and multimodal fusion. To address these issues, we propose a Fine-grained Multimodal Fusion Network (FMFN) for MSA. Firstly, a Fine-grained Representation Learning (FRL) module is designed to extract critical sentiment information from original modality features by employing finite learnable denoising tokens. FRL enables sentiment-related fine-grained representations to interact with these denoising tokens, thereby mitigating noise propagation and information redundancy. Secondly, a Token-level Cross-modal Alignment (TCA) module is introduced, which aligns different modality representations through fine-grained contrastive learning. TCA refines each modality by filtering out extraneous representations while facilitating subsequent multimodal interactions. Finally, a Correlation-aware Multimodal Fusion (CMF) module is developed to capture latent cross-modal correlations, generating consistent and complementary multimodal representations that enhance sentiment prediction. Extensive experiments are conducted on two English datasets, CMU-MOSI and CMU-MOSEI (aligned and unaligned versions), and one Chinese dataset, CH-SIMS (unaligned). The proposed method outperforms state-of-the-art models with 1%–2% accuracy and F1 score improvements, particularly on unaligned datasets, demonstrating the effectiveness of learning the fine-grained representation and token-level alignment in enhancing MSA performance.
AB - Multimodal sentiment analysis (MSA) has gained significant attention recently due to its crucial applications in intelligent systems. The effective extraction and fusion of sentiment representations from heterogeneous modalities (visual, text, and audio) remain the key and challenge in improving MSA performance. However, most existing methods often directly integrate different modalities and fail to capture fine-grained multimodal representations, suffering from irrelevant information among heterogeneous modalities. Moreover, these methods overlook the explicit token-level alignment between different modalities, which hinders effective representation learning and multimodal fusion. To address these issues, we propose a Fine-grained Multimodal Fusion Network (FMFN) for MSA. Firstly, a Fine-grained Representation Learning (FRL) module is designed to extract critical sentiment information from original modality features by employing finite learnable denoising tokens. FRL enables sentiment-related fine-grained representations to interact with these denoising tokens, thereby mitigating noise propagation and information redundancy. Secondly, a Token-level Cross-modal Alignment (TCA) module is introduced, which aligns different modality representations through fine-grained contrastive learning. TCA refines each modality by filtering out extraneous representations while facilitating subsequent multimodal interactions. Finally, a Correlation-aware Multimodal Fusion (CMF) module is developed to capture latent cross-modal correlations, generating consistent and complementary multimodal representations that enhance sentiment prediction. Extensive experiments are conducted on two English datasets, CMU-MOSI and CMU-MOSEI (aligned and unaligned versions), and one Chinese dataset, CH-SIMS (unaligned). The proposed method outperforms state-of-the-art models with 1%–2% accuracy and F1 score improvements, particularly on unaligned datasets, demonstrating the effectiveness of learning the fine-grained representation and token-level alignment in enhancing MSA performance.
KW - Fine-grained representation
KW - Multimodal fusion
KW - Representation learning
KW - Sentiment analysis
KW - Token-level alignment
UR - https://www.scopus.com/pages/publications/85214336033
U2 - 10.1016/j.eswa.2024.126274
DO - 10.1016/j.eswa.2024.126274
M3 - 文章
AN - SCOPUS:85214336033
SN - 0957-4174
VL - 269
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 126274
ER -