Skip to main navigation Skip to search Skip to main content

Pre-Segmentation: A Data Augmentation-Based Strategy for Thinning ViT Models

  • School of Computer Science and Technology, Anhui University
  • USTC-DEQING Alpha Innovation Institute

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In the basic Vision transformer (ViT) model, the way to process the image is to cut the image into a certain number of patches as labeled (location coded) tokens, which are inputted to the multi-attention part of the model to extract the information, and then passed through the full connectivity layer to a classifier to classify the image. If all the Tokens are utilized for interaction, the complex ViT structure brings a huge amount of computation, which is a major drawback of the model. The ViT model itself is characterized by global interaction, and Attention is extremely sensitive to global information, so the over-pursuit of interaction with all the Tokens will bring about a sharp increase in complexity and a multiplication of the pressure on the hardware facilities, such as the reduction in the number of Tokens, which will bring about a reduction in the amount of information. If the number of Tokens is reduced, it brings about loss of information and unsatisfactory accuracy. In this study, we take the lightweight ViT model as a starting point, and preprocess the images: the images are chunked and then entered into the coding layer for information extraction respectively, and finally the respective cls labels are processed to get the classification results, which can solve the problem of over-attention to global attention of ViT as well as avoiding the loss of information due to the reduction of Tokens, and ultimately realizing the model's lightweighting. In conclusion, our model (PS-ViT) starts from two aspects, firstly, data enhancement is used to improve the generalization ability of the model; secondly, in the processing of input images, pre-blocking can both change the range of model information interaction and reduce the computation of the model in order to achieve efficient reasoning. Finally, a 40% reduction in computation can be achieved while maintaining accuracy.

Original languageEnglish
Title of host publicationProceedings - 2024 10th International Conference on Big Data Computing and Communications, BIGCOM 2024
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages96-102
Number of pages7
Edition2024
ISBN (Electronic)9798331509538
DOIs
StatePublished - 2024
Event10th International Conference on Big Data Computing and Communications, BIGCOM 2024 - Dalian, China
Duration: 9 Aug 202411 Aug 2024

Conference

Conference10th International Conference on Big Data Computing and Communications, BIGCOM 2024
Country/TerritoryChina
CityDalian
Period9/08/2411/08/24

Fingerprint

Dive into the research topics of 'Pre-Segmentation: A Data Augmentation-Based Strategy for Thinning ViT Models'. Together they form a unique fingerprint.

Cite this