跳到主要导航 跳到搜索 跳到主要内容

Retraining-free Model Quantization via One-Shot Weight-Coupling Learning

  • Chen Tang
  • , Yuan Meng
  • , Jiacheng Jiang
  • , Shuzhao Xie
  • , Rongwei Lu
  • , Xinzhu Ma
  • , Zhi Wang*
  • , Wenwu Zhu*
  • *此作品的通讯作者
  • Tsinghua University
  • Chinese University of Hong Kong

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Quantization is of significance for compressing the over-parameterized deep neural models and deploying them on resource-limited devices. Fixed-precision quantization suffers from performance drop due to the limited numerical representation ability. Conversely, mixed-precision quantization (MPQ) is advocated to compress the model effectively by allocating heterogeneous bit-width for layers. MPQ is typically organized into a searching-retraining two-stage process. Previous works only focus on determining the optimal bit-width configuration in the first stage efficiently, while ignoring the considerable time costs in the second stage and thus hindering deployment efficiency significantly. In this paper, we devise a one-shot training-searching paradigm for mixed-precision model compression. Specifically, in the first stage, all potential bit-width configurations are coupled and thus optimized simultaneously within a set of shared weights. However, our observations reveal a previously unseen and severe bit-width interference phenomenon among highly coupled weights during optimization, leading to considerable performance degradation under a high compression ratio. To tackle this problem, we first design a bit-width scheduler to dynamically freeze the most turbulent bit-width of layers during training, to ensure the rest bit-widths converged properly. Then, taking inspiration from information theory, we present an information distortion mitigation technique to align the behaviour of the bad-performing bit-widths to the well-performing ones. In the second stage, an inference-only greedy search scheme is devised to evaluate the goodness of configurations without introducing any additional training costs. Extensive experiments on three representative models and three datasets demonstrate the effectiveness of the proposed method. Code can be available on https://github.com/1hunters/retraining-free-quantization.

源语言英语
主期刊名Proceedings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
出版商IEEE Computer Society
15855-15865
页数11
ISBN(电子版)9798350353006
ISBN(印刷版)9798350353006
DOI
出版状态已出版 - 2024
已对外发布
活动2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Seattle, 美国
期限: 16 6月 202422 6月 2024

出版系列

姓名Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
ISSN(印刷版)1063-6919

会议

会议2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
国家/地区美国
Seattle
时期16/06/2422/06/24

学术指纹

探究 'Retraining-free Model Quantization via One-Shot Weight-Coupling Learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此