Skip to main navigation Skip to search Skip to main content

Efficient Locality-aware Instruction Stream Scheduling for Stencil Computation on ARM Processors

  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Stencil computation is one of the fundamental computational patterns in scientific computing, commonly adopted in solving partial differential equations (PDEs) and a wide range of application fields. However, due to the memory-bound nature, it is challenging to achieve satisfactory performance on the ARM many-core processors with complex computation and memory hierarchies. In this study, we propose independent instruction stream scheduling with the Serial-FMA to Tree-Based Reduction (SFTBR) technique to decompose the stencil computation into multiple independent instruction streams for improved instruction-level parallelism. Furthermore, we propose a locality-aware block scheduling technique for locality-aware multi-level thread parallelism to address the complexities of cache and memory hierarchies on modern ARM many-core processors. Based on the above techniques, we implement a domain-specific compiler, AOStencil, to automatically generate optimized stencil codes on ARM many-core processors with genetic-algorithm-driven parameter tuning. Our evaluation results demonstrate that AOStencil achieves up to 4.39 × speedup over the state-of-the-art domain-specific compilers on Kunpeng and Phytium platforms.

Original languageEnglish
Title of host publicationACM ICS 2025 - Proceedings of the 39th ACM International Conference on Supercomputing
PublisherAssociation for Computing Machinery
Pages250-264
Number of pages15
ISBN (Electronic)9798400715372
DOIs
StatePublished - 22 Aug 2025
Event39th ACM International Conference on Supercomputing, ICS 2025 - Lake City, United States
Duration: 8 Jun 202511 Jun 2025

Publication series

NameProceedings of the International Conference on Supercomputing
VolumePart of 213821

Conference

Conference39th ACM International Conference on Supercomputing, ICS 2025
Country/TerritoryUnited States
CityLake City
Period8/06/2511/06/25

Keywords

  • Domain Specific Language
  • Many-core Architecture
  • Optimization Strategies
  • Stencil Computation

Fingerprint

Dive into the research topics of 'Efficient Locality-aware Instruction Stream Scheduling for Stencil Computation on ARM Processors'. Together they form a unique fingerprint.

Cite this