Abstract
In the basic Vision transformer (ViT) model, the way to process the image is to cut the image into a certain number of patches as labeled (location coded) tokens, which are inputted to the multi-attention part of the model to extract the information, and then passed through the full connectivity layer to a classifier to classify the image. If all the Tokens are utilized for interaction, the complex ViT structure brings a huge amount of computation, which is a major drawback of the model. The ViT model itself is characterized by global interaction, and Attention is extremely sensitive to global information, so the over-pursuit of interaction with all the Tokens will bring about a sharp increase in complexity and a multiplication of the pressure on the hardware facilities, such as the reduction in the number of Tokens, which will bring about a reduction in the amount of information. If the number of Tokens is reduced, it brings about loss of information and unsatisfactory accuracy. In this study, we take the lightweight ViT model as a starting point, and preprocess the images: the images are chunked and then entered into the coding layer for information extraction respectively, and finally the respective cls labels are processed to get the classification results, which can solve the problem of over-attention to global attention of ViT as well as avoiding the loss of information due to the reduction of Tokens, and ultimately realizing the model's lightweighting. In conclusion, our model (PS-ViT) starts from two aspects, firstly, data enhancement is used to improve the generalization ability of the model; secondly, in the processing of input images, pre-blocking can both change the range of model information interaction and reduce the computation of the model in order to achieve efficient reasoning. Finally, a 40% reduction in computation can be achieved while maintaining accuracy.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2024 10th International Conference on Big Data Computing and Communications, BIGCOM 2024 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 96-102 |
| Number of pages | 7 |
| Edition | 2024 |
| ISBN (Electronic) | 9798331509538 |
| DOIs | |
| State | Published - 2024 |
| Event | 10th International Conference on Big Data Computing and Communications, BIGCOM 2024 - Dalian, China Duration: 9 Aug 2024 → 11 Aug 2024 |
Conference
| Conference | 10th International Conference on Big Data Computing and Communications, BIGCOM 2024 |
|---|---|
| Country/Territory | China |
| City | Dalian |
| Period | 9/08/24 → 11/08/24 |
Fingerprint
Dive into the research topics of 'Pre-Segmentation: A Data Augmentation-Based Strategy for Thinning ViT Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver