跳到主要导航 跳到搜索 跳到主要内容

Distilling Cross-Modal Knowledge via Feature Disentanglement

  • Junhong Liu
  • , Yuan Zhang
  • , Tao Huang
  • , Wenchao Xu
  • , Renyu Yang*
  • *此作品的通讯作者
  • Beihang University
  • Peking University
  • Shanghai Jiao Tong University
  • Hong Kong University of Science and Technology

科研成果: 期刊稿件会议文章同行评审

摘要

Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-modal scenarios, such as vision-to-language distillation, where inconsistencies in representation across modalities lead to difficult knowledge transfer. To address this challenge, we propose frequency-decoupled cross-modal knowledge distillation, a method designed to decouple and balance knowledge transfer across modalities by leveraging frequency-domain features. We observed that low-frequency features exhibit high consistency across different modalities, whereas high-frequency features demonstrate extremely low cross-modal similarity. Accordingly, we apply distinct losses to these features: enforcing strong alignment in the low-frequency domain and introducing relaxed alignment for high-frequency features. We also propose a scale consistency loss to address distributional shifts between modalities, and employ a shared classifier to unify feature spaces. Extensive experiments across multiple benchmark datasets show our method substantially outperforms traditional KD and stateof-the-art cross-modal KD approaches.

源语言英语
页(从-至)23739-23747
页数9
期刊Proceedings of the AAAI Conference on Artificial Intelligence
40
28
DOI
出版状态已出版 - 2026
活动40th AAAI Conference on Artificial Intelligence, AAAI 2026 - Singapore, 新加坡
期限: 20 1月 202627 1月 2026

指纹

探究 'Distilling Cross-Modal Knowledge via Feature Disentanglement' 的科研主题。它们共同构成独一无二的指纹。

引用此