Skip to main navigation Skip to search Skip to main content

MixMixQ: Quantization with Mixed Bit-Sparsity and Mixed Bit-Width for CIM Accelerators

  • Jinyu Bai
  • , He Zhang
  • , Long Chao Liu
  • , Pengfei Li
  • , Wang Kang*
  • *Corresponding author for this work
  • Beihang University
  • Beijing Jinghanyu Electronic Engineering Technology Co. Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Quantization is vital for deploying neural networks on Computing-In-Memory (CIM) based accelerators due to inherent limitations in memory devices and data interfaces' representational capacities. However, traditional quantization algorithms often overlook CIM's unique computing paradigm, leading to suboptimal performance. To address this, we introduce MixMixQ, a novel quantization algorithm specifically designed for CIM accelerators that strategically integrates mixed bit-sparsity and mixed bit-width, enhancing overall hardware efficiency while preserving high accuracy. Notably, our method can enhance hardware efficiency by up to 294% compared to traditional quantization methods, with only a minimal 0.13% decrease in accuracy compared to a full-precision network.

Original languageEnglish
Title of host publicationGLSVLSI 2024 - Proceedings of the Great Lakes Symposium on VLSI 2024
PublisherAssociation for Computing Machinery
Pages537-540
Number of pages4
ISBN (Electronic)9798400706059
DOIs
StatePublished - 12 Jun 2024
Event34th Great Lakes Symposium on VLSI 2024, GLSVLSI 2024 - Clearwater, United States
Duration: 12 Jun 202414 Jun 2024

Publication series

NameProceedings of the ACM Great Lakes Symposium on VLSI, GLSVLSI

Conference

Conference34th Great Lakes Symposium on VLSI 2024, GLSVLSI 2024
Country/TerritoryUnited States
CityClearwater
Period12/06/2414/06/24

Keywords

  • Bit-Sparsity
  • Computing-In-Memory
  • Evolutionary Algorithm
  • Gradient Estimation
  • Neural Network Quantization

Fingerprint

Dive into the research topics of 'MixMixQ: Quantization with Mixed Bit-Sparsity and Mixed Bit-Width for CIM Accelerators'. Together they form a unique fingerprint.

Cite this