Abstract
Sparse generalized matrix multiplication (SpGEMM) has been widely applied to sparse neural network models. However, the arbitrary distribution of non-zero elements in sparse matrices leads to the processing elements (PEs) in the systolic array (SA) architecture being idle and further affecting computing efficiency. Reviewing existing methods, we found three main drawbacks to calculating SpGEMM in multi-core SAs. First, the sparse matrix calculation format is unsuitable for the SA architecture. Second, when the SA calculates SpGEMM, the load is unbalanced among PEs. Third, the computational load is unevenly distributed across different SA cores during the process above. To address the above problems, we proposed a load-balancing SpGEMM accelerating framework for multi-core SAs. First, we introduced the SCSR sparse matrix compression format and the PE fast sparse matrix matching and calculation method in SA. Second, we present a runtime dynamic data flow packaging algorithm, GrePack. Third, we propose a compile-time sparse data flow multi-core static partitioning algorithm. Compared with the state-of-the-art work, our dynamic packaging algorithm accelerates SpGEMM speed by up to 2.08×, our static multi-core partitioning method improves the computing unit utilization by up to 1.54×, and our collaborative inference framework improves SpGEMM speed by up to 2.29×.
| Original language | English |
|---|---|
| Article number | 103186 |
| Journal | Parallel Computing |
| Volume | 127 |
| DOIs | |
| State | Published - Mar 2026 |
Keywords
- Hardware acceleration
- Load-balancing dataflow
- Multi-core
- SpGEMM
- Sparse systolic array
Fingerprint
Dive into the research topics of 'LSAF: A load-balancing SpGEMM acceleration framework with dynamic package and static partition for multi-core systolic arrays'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver