TY - GEN
T1 - Toward accelerated stencil computation by adapting tensor core unit on GPU
AU - Liu, Xiaoyan
AU - Liu, Yi
AU - Yang, Hailong
AU - Liao, Jianjin
AU - Li, Mingzhen
AU - Luan, Zhongzhi
AU - Qian, Depei
N1 - Publisher Copyright:
© 2022 ACM.
PY - 2022/6/28
Y1 - 2022/6/28
N2 - The Tensor Core Unit (TCU) has been increasingly adopted on modern high performance processors, specialized in boosting the performance of general matrix multiplication (GEMM). Due to its highly optimized hardware design, TCU can significantly accelerate GEMM-based operations widely used in scientific as well as deep learning applications. However, there is few work exploiting TCU to accelerate non-GEMM operations such as stencil computation that is also important in the field of high performance computing. To the best of our knowledge, there is no previous work that adapts stencil computation to TCU efficiently by considering its unique characteristics. In this paper, we propose a new method called TCstencil to adapt TCU for accelerating stencil computation. Specifically, we re-design the stencil computation as a series of reduction and summation operations in order to leverage the computing power of TCU. In addition, we propose corresponding optimizations for better exploiting TCU and memory hierarchy on GPU. We evaluate our method with different stencils and input mesh sizes on NVIDIA A100 and V100 GPUs. The experiment results demonstrate our method can achieve superior performance compared to the state-of-the-art stencil optimization frameworks.
AB - The Tensor Core Unit (TCU) has been increasingly adopted on modern high performance processors, specialized in boosting the performance of general matrix multiplication (GEMM). Due to its highly optimized hardware design, TCU can significantly accelerate GEMM-based operations widely used in scientific as well as deep learning applications. However, there is few work exploiting TCU to accelerate non-GEMM operations such as stencil computation that is also important in the field of high performance computing. To the best of our knowledge, there is no previous work that adapts stencil computation to TCU efficiently by considering its unique characteristics. In this paper, we propose a new method called TCstencil to adapt TCU for accelerating stencil computation. Specifically, we re-design the stencil computation as a series of reduction and summation operations in order to leverage the computing power of TCU. In addition, we propose corresponding optimizations for better exploiting TCU and memory hierarchy on GPU. We evaluate our method with different stencils and input mesh sizes on NVIDIA A100 and V100 GPUs. The experiment results demonstrate our method can achieve superior performance compared to the state-of-the-art stencil optimization frameworks.
KW - GPU
KW - Performance optimization
KW - Stencil computation
KW - Tensor core
UR - https://www.scopus.com/pages/publications/85132841998
U2 - 10.1145/3524059.3532392
DO - 10.1145/3524059.3532392
M3 - 会议稿件
AN - SCOPUS:85132841998
T3 - Proceedings of the International Conference on Supercomputing
BT - Proceedings of the 36th ACM International Conference on Supercomputing, ICS 2022
PB - Association for Computing Machinery
T2 - 36th ACM International Conference on Supercomputing, ICS 2022
Y2 - 27 June 2022 through 30 June 2022
ER -