跳到主要导航 跳到搜索 跳到主要内容

Self-slimmed Vision Transformer

  • Zhuofan Zong
  • , Kunchang Li
  • , Guanglu Song
  • , Yali Wang
  • , Yu Qiao
  • , Biao Leng
  • , Yu Liu*
  • *此作品的通讯作者
  • Beihang University
  • Shenzhen Institute of Advanced Technology
  • University of Chinese Academy of Sciences
  • SenseTime Group Limited
  • Shenzhen Institute of Artificial Intelligence and Robotics for Society
  • Shanghai Artificial Intelligence Laboratory

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Vision transformers (ViTs) have become the popular structures and outperformed convolutional neural networks (CNNs) on various vision tasks. However, such powerful transformers bring a huge computation burden, because of the exhausting token-to-token comparison. The previous works focus on dropping insignificant tokens to reduce the computational cost of ViTs. But when the dropping ratio increases, this hard manner will inevitably discard the vital tokens, which limits its efficiency. To solve the issue, we propose a generic self-slimmed learning approach for vanilla ViTs, namely SiT. Specifically, we first design a novel Token Slimming Module (TSM), which can boost the inference efficiency of ViTs by dynamic token aggregation. As a general method of token hard dropping, our TSM softly integrates redundant tokens into fewer informative ones. It can dynamically zoom visual attention without cutting off discriminative token relations in the images, even with a high slimming ratio. Furthermore, we introduce a concise Feature Recal-ibration Distillation (FRD) framework, wherein we design a reverse version of TSM (RTSM) to recalibrate the unstructured token in a flexible auto-encoder manner. Due to the similar structure between teacher and student, our FRD can effectively leverage structure knowledge for better convergence. Finally, we conduct extensive experiments to evaluate our SiT. It demonstrates that our method can speed up ViTs by 1.7× with negligible accuracy drop, and even speed up ViTs by 3.6× while maintaining 97% of their performance. Surprisingly, by simply arming LV-ViT with our SiT, we achieve new state-of-the-art performance on ImageNet. Code is available at https://github.com/Sense-X/SiT.

源语言英语
主期刊名Computer Vision – ECCV 2022 - 17th European Conference, Proceedings
编辑Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, Tal Hassner
出版商Springer Science and Business Media Deutschland GmbH
432-448
页数17
ISBN(印刷版)9783031200823
DOI
出版状态已出版 - 2022
活动17th European Conference on Computer Vision, ECCV 2022 - Tel Aviv, 以色列
期限: 23 10月 202227 10月 2022

出版系列

姓名Lecture Notes in Computer Science
13671 LNCS
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议17th European Conference on Computer Vision, ECCV 2022
国家/地区以色列
Tel Aviv
时期23/10/2227/10/22

学术指纹

探究 'Self-slimmed Vision Transformer' 的科研主题。它们共同构成独一无二的学术指纹。

引用此