TY - JOUR
T1 - An End-to-End In-Memory Computing System Based on a 40-nm eFlash-Based IMC SoC
T2 - Circuits, Toolchains, and Systems Co-Design Framework
AU - Bai, Tianshuo
AU - Mao, Wanru
AU - Wang, Guangyao
AU - Liu, Hanjie
AU - Zhang, Aifei
AU - Fu, Shihang
AU - Liu, Shuaikai
AU - Hu, Jianchao
AU - Yang, Xitong
AU - Pan, Biao
AU - Xing, Wei
AU - Kang, Wang
N1 - Publisher Copyright:
© 1982-2012 IEEE.
PY - 2024/6/1
Y1 - 2024/6/1
N2 - Despite its promising potential for artificial intelligence (AI) applications, current in-memory computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance, including 1) an 8-bit hardware-friendly quantization-aware training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data; 2) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips; and 3) an efficient mapping strategy based on the integer linear programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40-nm embedded Flash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 h for voice recognition, a 21.53% improvement for the perceptual evaluation of speech quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection.
AB - Despite its promising potential for artificial intelligence (AI) applications, current in-memory computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance, including 1) an 8-bit hardware-friendly quantization-aware training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data; 2) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips; and 3) an efficient mapping strategy based on the integer linear programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40-nm embedded Flash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 h for voice recognition, a 21.53% improvement for the perceptual evaluation of speech quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection.
KW - Circuit-toolchain-system co-design framework
KW - embedded Flash (eFlash)
KW - in-memory computing (IMC)
KW - toolchains
UR - https://www.scopus.com/pages/publications/85182346732
U2 - 10.1109/TCAD.2024.3349502
DO - 10.1109/TCAD.2024.3349502
M3 - 文章
AN - SCOPUS:85182346732
SN - 0278-0070
VL - 43
SP - 1729
EP - 1740
JO - IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
JF - IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
IS - 6
ER -