摘要
In response to the security risks posed by the realistic propagation of manipulated media data, detecting and grounding multi-modal media manipulation has received attention as a challenging task. However, there is multi-modal contribution imbalance on current approach for cross-modal learning, which affects model performance optimisation. To this end, we propose an Adaptive Contribution Modulation (ACM) framework to solve the problem of multi-modal contribution imbalance. To balance the image and text embedding features before fusion, we propose adaptive weight decision to computes dynamic weights for fusion features, which enable more adaptive and robust decision-making. Meanwhile, we propose contribution modulation block, which dynamically governs the contributions of different modalities for optimization. Based on cross-modal contrastive learning, we balance image and text embeddings contribution through multi-modal contribution balanced learning, which makes better use of the semantic correlation of all modalities. We conduct experiments on the DGM4 dataset, which demonstrate the superior performance of our approach through compared to state-of-the-art methods.
| 源语言 | 英语 |
|---|---|
| 期刊 | Proceedings - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing |
| DOI | |
| 出版状态 | 已出版 - 2025 |
| 活动 | 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Hyderabad, 印度 期限: 6 4月 2025 → 11 4月 2025 |
学术指纹
探究 'Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver