TY - GEN
T1 - A Computing Efficient Hardware Architecture for Sparse Deep Neural Network Computing
AU - Zhang, Yanwen
AU - Ouyang, Peng
AU - Yin, Shouyi
AU - Zhang, Youguang
AU - Zhao, Weisheng
AU - Wei, Shaojun
N1 - Publisher Copyright:
© 2018 IEEE.
PY - 2018/12/5
Y1 - 2018/12/5
N2 - Convolutional Neural Networks (CNNs) have demonstrated significant performance in AI (artificial intelligence) systems. However, CNNs often have tens or even hundreds of neural layers with millions of parameters to achieve state-of-the-art performance, which hinders the deployment to some resource limited scenarios. Meanwhile, those parameters and data usually are sparse, which results in useless calculation as well as unbalanced calculation. To solve these problem, we propose a computing efficient hardware architecture. In order to decrease calculating redundancy, we filter zero-valued weights and zero-valued feature maps. To reduce redundant memory consumption, we propose a memory division and a data reuse mechanism. To resolve load imbalance, we implement a near-zero-cost scheduling switching strategy. Experimental results show that our architecture saves, on average, 22.6% memory times and 60.5% computing time over the state-of-the-art NN accelerator.
AB - Convolutional Neural Networks (CNNs) have demonstrated significant performance in AI (artificial intelligence) systems. However, CNNs often have tens or even hundreds of neural layers with millions of parameters to achieve state-of-the-art performance, which hinders the deployment to some resource limited scenarios. Meanwhile, those parameters and data usually are sparse, which results in useless calculation as well as unbalanced calculation. To solve these problem, we propose a computing efficient hardware architecture. In order to decrease calculating redundancy, we filter zero-valued weights and zero-valued feature maps. To reduce redundant memory consumption, we propose a memory division and a data reuse mechanism. To resolve load imbalance, we implement a near-zero-cost scheduling switching strategy. Experimental results show that our architecture saves, on average, 22.6% memory times and 60.5% computing time over the state-of-the-art NN accelerator.
KW - CNN
KW - Dataflow processing
KW - Spatial Architecture
UR - https://www.scopus.com/pages/publications/85060289717
U2 - 10.1109/ICSICT.2018.8565755
DO - 10.1109/ICSICT.2018.8565755
M3 - 会议稿件
AN - SCOPUS:85060289717
T3 - 2018 14th IEEE International Conference on Solid-State and Integrated Circuit Technology, ICSICT 2018 - Proceedings
BT - 2018 14th IEEE International Conference on Solid-State and Integrated Circuit Technology, ICSICT 2018 - Proceedings
A2 - Tang, Ting-Ao
A2 - Ye, Fan
A2 - Jiang, Yu-Long
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 14th IEEE International Conference on Solid-State and Integrated Circuit Technology, ICSICT 2018
Y2 - 31 October 2018 through 3 November 2018
ER -