跳到主要导航 跳到搜索 跳到主要内容

Toward accelerated stencil computation by adapting tensor core unit on GPU

  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The Tensor Core Unit (TCU) has been increasingly adopted on modern high performance processors, specialized in boosting the performance of general matrix multiplication (GEMM). Due to its highly optimized hardware design, TCU can significantly accelerate GEMM-based operations widely used in scientific as well as deep learning applications. However, there is few work exploiting TCU to accelerate non-GEMM operations such as stencil computation that is also important in the field of high performance computing. To the best of our knowledge, there is no previous work that adapts stencil computation to TCU efficiently by considering its unique characteristics. In this paper, we propose a new method called TCstencil to adapt TCU for accelerating stencil computation. Specifically, we re-design the stencil computation as a series of reduction and summation operations in order to leverage the computing power of TCU. In addition, we propose corresponding optimizations for better exploiting TCU and memory hierarchy on GPU. We evaluate our method with different stencils and input mesh sizes on NVIDIA A100 and V100 GPUs. The experiment results demonstrate our method can achieve superior performance compared to the state-of-the-art stencil optimization frameworks.

源语言英语
主期刊名Proceedings of the 36th ACM International Conference on Supercomputing, ICS 2022
出版商Association for Computing Machinery
ISBN(电子版)9781450392815
DOI
出版状态已出版 - 28 6月 2022
活动36th ACM International Conference on Supercomputing, ICS 2022 - Virtual, Online
期限: 27 6月 202230 6月 2022

丛书

姓名Proceedings of the International Conference on Supercomputing

会议

会议36th ACM International Conference on Supercomputing, ICS 2022
Virtual, Online
时期27/06/2230/06/22

学术指纹

探究 'Toward accelerated stencil computation by adapting tensor core unit on GPU' 的科研主题。它们共同构成独一无二的学术指纹。

引用此