TY - GEN
T1 - HACompBench
T2 - 25th International Conference on Algorithms and Architectures for Parallel Processing, ICA3PP 2025
AU - Gan, Zhengyu
AU - Du, Haohua
AU - Feng, Chengquan
AU - Tan, Haisheng
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Deployment of deep neural networks on edge devices faces challenges from heterogeneous hardware and multimodal tasks, where existing compression evaluation frameworks overlook hardware co-design, leading to suboptimal performance. To address this, we introduce HACompBench, a new hardware-aware framework that defines compression evaluation as a multi-objective optimization problem and combines hardware metrics such as quantization efficiency ξ and sparsity compatibility η with a dynamic scoring function J. We performed comprehensive experiments across four leading SoC platforms: Snapdragon 888, Snapdragon 765G, Kirin 970, and Jetson Nano P3450, and tested ten DNN models covering vision, text, and speech modalities using compression techniques such as quantization, pruning, and weight sharing, revealing hardware-induced performance gaps, such as quantization yields J=32.3% on Snapdragon 888 but J=68.0% on Jetson Nano P3450 due to INT8 emulation overhead. These results systematically highlight compression variations from differences in parallel processing capabilities. HACompBench innovates by linking hardware features, supporting multimodal tasks, and surpassing MLPerf Tiny’s single-modality focus and AIoTBench’s lack of co-design through embedded metrics ξ and η, while modality-specific corrections improve accuracy by up to 12.6%. It provides a unified and robust framework for edge deployment.
AB - Deployment of deep neural networks on edge devices faces challenges from heterogeneous hardware and multimodal tasks, where existing compression evaluation frameworks overlook hardware co-design, leading to suboptimal performance. To address this, we introduce HACompBench, a new hardware-aware framework that defines compression evaluation as a multi-objective optimization problem and combines hardware metrics such as quantization efficiency ξ and sparsity compatibility η with a dynamic scoring function J. We performed comprehensive experiments across four leading SoC platforms: Snapdragon 888, Snapdragon 765G, Kirin 970, and Jetson Nano P3450, and tested ten DNN models covering vision, text, and speech modalities using compression techniques such as quantization, pruning, and weight sharing, revealing hardware-induced performance gaps, such as quantization yields J=32.3% on Snapdragon 888 but J=68.0% on Jetson Nano P3450 due to INT8 emulation overhead. These results systematically highlight compression variations from differences in parallel processing capabilities. HACompBench innovates by linking hardware features, supporting multimodal tasks, and surpassing MLPerf Tiny’s single-modality focus and AIoTBench’s lack of co-design through embedded metrics ξ and η, while modality-specific corrections improve accuracy by up to 12.6%. It provides a unified and robust framework for edge deployment.
KW - Edge Computing
KW - Hardware-Algorithm Co-Design
KW - Model Compression
KW - Multimodal Evaluation
KW - Parallel Sparse Optimization
UR - https://www.scopus.com/pages/publications/105035831136
U2 - 10.1007/978-981-95-8405-5_26
DO - 10.1007/978-981-95-8405-5_26
M3 - 会议稿件
AN - SCOPUS:105035831136
SN - 9789819584048
T3 - Lecture Notes in Computer Science
SP - 478
EP - 496
BT - Algorithms and Architectures for Parallel Processing - 25th International Conference, ICA3PP 2025, Proceedings
A2 - Liu, Huazhong
A2 - Ibrahim, Shadi
A2 - Rauber, Thomas
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 30 October 2025 through 2 November 2025
ER -