跳到主要导航 跳到搜索 跳到主要内容

An empirical study of Qwen3 quantization

  • Xingyu Zheng
  • , Yuye Li
  • , Haoran Chu
  • , Yue Feng
  • , Xudong Ma
  • , Zining Wang
  • , Jie Luo
  • , Jinyang Guo
  • , Haotong Qin*
  • , Michele Magno
  • , Xianglong Liu
  • *此作品的通讯作者
  • Beihang University
  • Xidian University
  • ByteDance Ltd.
  • ETH Zurich - Institute for Particle Physics and Astrophysics (IPA)

科研成果: 期刊稿件文章同行评审

摘要

The Qwen series has emerged as a leading family of open-source large language models (LLMs), demonstrating remarkable capabilities in natural language understanding tasks. With the recent release of Qwen3, which exhibits superior performance across diverse benchmarks, there is an increased interest in the efficient deployment of these models in resource-constrained environments. Low-bit quantization presents a promising solution, yet its impact on Qwen3’s performance remains underexplored. This study conducts a systematic evaluation of Qwen3’s robustness under various quantization settings, aiming to identify both the opportunities and the challenges inherent in compressing this state-of-the-art model. We rigorously assess 5 existing classic post-training quantization techniques applied to Qwen3, spanning bit-widths from 1 to 8 bits, and evaluate their effectiveness across multiple datasets. Our findings reveal that while Qwen3 maintains competitive performance at moderate bit-widths, it experiences notable degradation in linguistic tasks under ultra-low precision, underscoring the persistent hurdles in LLM compression. These results emphasize the need for further research to mitigate performance loss in extreme quantization scenarios. We anticipate that this empirical analysis will provide actionable insights for advancing quantization methods tailored to Qwen3 and future LLMs, ultimately enhancing their practicality without compromising accuracy.

源语言英语
文章编号11
期刊Visual Intelligence
4
1
DOI
出版状态已出版 - 12月 2026

指纹

探究 'An empirical study of Qwen3 quantization' 的科研主题。它们共同构成独一无二的指纹。

引用此