跳到主要导航 跳到搜索 跳到主要内容

Efficient Locality-aware Instruction Stream Scheduling for Stencil Computation on ARM Processors

  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Stencil computation is one of the fundamental computational patterns in scientific computing, commonly adopted in solving partial differential equations (PDEs) and a wide range of application fields. However, due to the memory-bound nature, it is challenging to achieve satisfactory performance on the ARM many-core processors with complex computation and memory hierarchies. In this study, we propose independent instruction stream scheduling with the Serial-FMA to Tree-Based Reduction (SFTBR) technique to decompose the stencil computation into multiple independent instruction streams for improved instruction-level parallelism. Furthermore, we propose a locality-aware block scheduling technique for locality-aware multi-level thread parallelism to address the complexities of cache and memory hierarchies on modern ARM many-core processors. Based on the above techniques, we implement a domain-specific compiler, AOStencil, to automatically generate optimized stencil codes on ARM many-core processors with genetic-algorithm-driven parameter tuning. Our evaluation results demonstrate that AOStencil achieves up to 4.39 × speedup over the state-of-the-art domain-specific compilers on Kunpeng and Phytium platforms.

源语言英语
主期刊名ACM ICS 2025 - Proceedings of the 39th ACM International Conference on Supercomputing
出版商Association for Computing Machinery
250-264
页数15
ISBN(电子版)9798400715372
DOI
出版状态已出版 - 22 8月 2025
活动39th ACM International Conference on Supercomputing, ICS 2025 - Lake City, 美国
期限: 8 6月 202511 6月 2025

丛书

姓名Proceedings of the International Conference on Supercomputing
Part of 213821

会议

会议39th ACM International Conference on Supercomputing, ICS 2025
国家/地区美国
Lake City
时期8/06/2511/06/25

学术指纹

探究 'Efficient Locality-aware Instruction Stream Scheduling for Stencil Computation on ARM Processors' 的科研主题。它们共同构成独一无二的学术指纹。

引用此