跳到主要导航 跳到搜索 跳到主要内容

Energy-Efficient Parallel Ising SoftMax Engine for Transformer Attention Based on Asynchronous SOT-P-Bits

  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Transformer models have achieved remarkable progress in both natural language processing and computer vision; however, the SoftMax operation has become a major bottleneck in Transformer inference due to costly exponential operations and repeated memory access. We present a hardware-native SoftMax engine based on an asynchronous probabilistic Ising machine built from experimentally fabricated SOT-MTJ P-Bit. A SoftMax-To-Ising compiler maps arbitrary logits into an implementable energy function, enabling direct physical sampling of the corresponding Boltzmann distribution. The resulting Ising SoftMax converges rapidly, exceeding FP16 precision at 32μ text {s}. When integrated into neural network workloads, it preserves high inference accuracy without noticeable degradation. These results demonstrate a fully asynchronous physical SoftMax engine and highlight its potential for energy-efficient deep learning acceleration.

源语言英语
页(从-至)1458-1461
页数4
期刊IEEE Electron Device Letters
47
7
DOI
出版状态已出版 - 1 7月 2026

学术指纹

探究 'Energy-Efficient Parallel Ising SoftMax Engine for Transformer Attention Based on Asynchronous SOT-P-Bits' 的科研主题。它们共同构成独一无二的学术指纹。

引用此