摘要
Transformer models have achieved remarkable progress in both natural language processing and computer vision; however, the SoftMax operation has become a major bottleneck in Transformer inference due to costly exponential operations and repeated memory access. We present a hardware-native SoftMax engine based on an asynchronous probabilistic Ising machine built from experimentally fabricated SOT-MTJ P-Bit. A SoftMax-To-Ising compiler maps arbitrary logits into an implementable energy function, enabling direct physical sampling of the corresponding Boltzmann distribution. The resulting Ising SoftMax converges rapidly, exceeding FP16 precision at 32μ text {s}. When integrated into neural network workloads, it preserves high inference accuracy without noticeable degradation. These results demonstrate a fully asynchronous physical SoftMax engine and highlight its potential for energy-efficient deep learning acceleration.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 1458-1461 |
| 页数 | 4 |
| 期刊 | IEEE Electron Device Letters |
| 卷 | 47 |
| 期 | 7 |
| DOI | |
| 出版状态 | 已出版 - 1 7月 2026 |
学术指纹
探究 'Energy-Efficient Parallel Ising SoftMax Engine for Transformer Attention Based on Asynchronous SOT-P-Bits' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver