跳到主要导航 跳到搜索 跳到主要内容

Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding

  • Yixiang Li
  • , Biao Leng*
  • *此作品的通讯作者
  • Beihang University

科研成果: 期刊稿件会议文章同行评审

摘要

In response to the security risks posed by the realistic propagation of manipulated media data, detecting and grounding multi-modal media manipulation has received attention as a challenging task. However, there is multi-modal contribution imbalance on current approach for cross-modal learning, which affects model performance optimisation. To this end, we propose an Adaptive Contribution Modulation (ACM) framework to solve the problem of multi-modal contribution imbalance. To balance the image and text embedding features before fusion, we propose adaptive weight decision to computes dynamic weights for fusion features, which enable more adaptive and robust decision-making. Meanwhile, we propose contribution modulation block, which dynamically governs the contributions of different modalities for optimization. Based on cross-modal contrastive learning, we balance image and text embeddings contribution through multi-modal contribution balanced learning, which makes better use of the semantic correlation of all modalities. We conduct experiments on the DGM4 dataset, which demonstrate the superior performance of our approach through compared to state-of-the-art methods.

学术指纹

探究 'Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding' 的科研主题。它们共同构成独一无二的学术指纹。

引用此