Abstract
Transformer models have achieved remarkable progress in both natural language processing and computer vision; however, the SoftMax operation has become a major bottleneck in Transformer inference due to costly exponential operations and repeated memory access. We present a hardware-native SoftMax engine based on an asynchronous probabilistic Ising machine built from experimentally fabricated SOT-MTJ P-Bit. A SoftMax-To-Ising compiler maps arbitrary logits into an implementable energy function, enabling direct physical sampling of the corresponding Boltzmann distribution. The resulting Ising SoftMax converges rapidly, exceeding FP16 precision at 32μ text {s}. When integrated into neural network workloads, it preserves high inference accuracy without noticeable degradation. These results demonstrate a fully asynchronous physical SoftMax engine and highlight its potential for energy-efficient deep learning acceleration.
| Original language | English |
|---|---|
| Pages (from-to) | 1458-1461 |
| Number of pages | 4 |
| Journal | IEEE Electron Device Letters |
| Volume | 47 |
| Issue number | 7 |
| DOIs | |
| State | Published - 1 Jul 2026 |
Keywords
- Ising machine
- P-Bit
- SOT-MTJ
- transformer
Fingerprint
Dive into the research topics of 'Energy-Efficient Parallel Ising SoftMax Engine for Transformer Attention Based on Asynchronous SOT-P-Bits'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver