Skip to main navigation Skip to search Skip to main content

Toward accelerated stencil computation by adapting tensor core unit on GPU

  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The Tensor Core Unit (TCU) has been increasingly adopted on modern high performance processors, specialized in boosting the performance of general matrix multiplication (GEMM). Due to its highly optimized hardware design, TCU can significantly accelerate GEMM-based operations widely used in scientific as well as deep learning applications. However, there is few work exploiting TCU to accelerate non-GEMM operations such as stencil computation that is also important in the field of high performance computing. To the best of our knowledge, there is no previous work that adapts stencil computation to TCU efficiently by considering its unique characteristics. In this paper, we propose a new method called TCstencil to adapt TCU for accelerating stencil computation. Specifically, we re-design the stencil computation as a series of reduction and summation operations in order to leverage the computing power of TCU. In addition, we propose corresponding optimizations for better exploiting TCU and memory hierarchy on GPU. We evaluate our method with different stencils and input mesh sizes on NVIDIA A100 and V100 GPUs. The experiment results demonstrate our method can achieve superior performance compared to the state-of-the-art stencil optimization frameworks.

Original languageEnglish
Title of host publicationProceedings of the 36th ACM International Conference on Supercomputing, ICS 2022
PublisherAssociation for Computing Machinery
ISBN (Electronic)9781450392815
DOIs
StatePublished - 28 Jun 2022
Event36th ACM International Conference on Supercomputing, ICS 2022 - Virtual, Online
Duration: 27 Jun 202230 Jun 2022

Publication series

NameProceedings of the International Conference on Supercomputing

Conference

Conference36th ACM International Conference on Supercomputing, ICS 2022
CityVirtual, Online
Period27/06/2230/06/22

Keywords

  • GPU
  • Performance optimization
  • Stencil computation
  • Tensor core

Fingerprint

Dive into the research topics of 'Toward accelerated stencil computation by adapting tensor core unit on GPU'. Together they form a unique fingerprint.

Cite this