跳到主要导航 跳到搜索 跳到主要内容

Visformer: The Vision-friendly Transformer

  • Beihang University
  • Johns Hopkins University
  • Zhengzhou University
  • University of Science and Technology of China
  • Xidian University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The past year has witnessed the rapid development of applying the Transformer module to vision problems. While some researchers have demonstrated that Transformer-based models enjoy a favorable ability of fitting data, there are still growing number of evidences showing that these models suffer over-fitting especially when the training data is limited. This paper offers an empirical study by performing step-by-step operations to gradually transit a Transformer-based model to a convolution-based model. The results we obtain during the transition process deliver useful messages for improving visual recognition. Based on these observations, we propose a new architecture named Visformer, which is abbreviated from the 'Vision-friendly Transformer'. With the same computational complexity, Visformer outperforms both the Transformer-based and convolution-based models in terms of ImageNet classification accuracy, and the advantage becomes more significant when the model complexity is lower or the training set is smaller. The code is available at https://github.com/danczs/Visformer.

源语言英语
主期刊名Proceedings - 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021
出版商Institute of Electrical and Electronics Engineers Inc.
569-578
页数10
ISBN(电子版)9781665428125
DOI
出版状态已出版 - 2021
活动18th IEEE/CVF International Conference on Computer Vision, ICCV 2021 - Virtual, Online, 加拿大
期限: 11 10月 202117 10月 2021

出版系列

姓名Proceedings of the IEEE International Conference on Computer Vision
ISSN(印刷版)1550-5499

会议

会议18th IEEE/CVF International Conference on Computer Vision, ICCV 2021
国家/地区加拿大
Virtual, Online
时期11/10/2117/10/21

指纹

探究 'Visformer: The Vision-friendly Transformer' 的科研主题。它们共同构成独一无二的指纹。

引用此