TY - GEN
T1 - Efficient Locality-aware Instruction Stream Scheduling for Stencil Computation on ARM Processors
AU - Liu, Shanghao
AU - Yang, Hailong
AU - You, Xin
AU - Luan, Zhongzhi
AU - Liu, Yi
AU - Qian, Depei
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s). Publication rights licensed to ACM.
PY - 2025/8/22
Y1 - 2025/8/22
N2 - Stencil computation is one of the fundamental computational patterns in scientific computing, commonly adopted in solving partial differential equations (PDEs) and a wide range of application fields. However, due to the memory-bound nature, it is challenging to achieve satisfactory performance on the ARM many-core processors with complex computation and memory hierarchies. In this study, we propose independent instruction stream scheduling with the Serial-FMA to Tree-Based Reduction (SFTBR) technique to decompose the stencil computation into multiple independent instruction streams for improved instruction-level parallelism. Furthermore, we propose a locality-aware block scheduling technique for locality-aware multi-level thread parallelism to address the complexities of cache and memory hierarchies on modern ARM many-core processors. Based on the above techniques, we implement a domain-specific compiler, AOStencil, to automatically generate optimized stencil codes on ARM many-core processors with genetic-algorithm-driven parameter tuning. Our evaluation results demonstrate that AOStencil achieves up to 4.39 × speedup over the state-of-the-art domain-specific compilers on Kunpeng and Phytium platforms.
AB - Stencil computation is one of the fundamental computational patterns in scientific computing, commonly adopted in solving partial differential equations (PDEs) and a wide range of application fields. However, due to the memory-bound nature, it is challenging to achieve satisfactory performance on the ARM many-core processors with complex computation and memory hierarchies. In this study, we propose independent instruction stream scheduling with the Serial-FMA to Tree-Based Reduction (SFTBR) technique to decompose the stencil computation into multiple independent instruction streams for improved instruction-level parallelism. Furthermore, we propose a locality-aware block scheduling technique for locality-aware multi-level thread parallelism to address the complexities of cache and memory hierarchies on modern ARM many-core processors. Based on the above techniques, we implement a domain-specific compiler, AOStencil, to automatically generate optimized stencil codes on ARM many-core processors with genetic-algorithm-driven parameter tuning. Our evaluation results demonstrate that AOStencil achieves up to 4.39 × speedup over the state-of-the-art domain-specific compilers on Kunpeng and Phytium platforms.
KW - Domain Specific Language
KW - Many-core Architecture
KW - Optimization Strategies
KW - Stencil Computation
UR - https://www.scopus.com/pages/publications/105021479836
U2 - 10.1145/3721145.3725760
DO - 10.1145/3721145.3725760
M3 - 会议稿件
AN - SCOPUS:105021479836
T3 - Proceedings of the International Conference on Supercomputing
SP - 250
EP - 264
BT - ACM ICS 2025 - Proceedings of the 39th ACM International Conference on Supercomputing
PB - Association for Computing Machinery
T2 - 39th ACM International Conference on Supercomputing, ICS 2025
Y2 - 8 June 2025 through 11 June 2025
ER -