Skip to main navigation Skip to search Skip to main content

Energy-Efficient Parallel Ising SoftMax Engine for Transformer Attention Based on Asynchronous SOT-P-Bits

  • Beihang University

Research output: Contribution to journalArticlepeer-review

Abstract

Transformer models have achieved remarkable progress in both natural language processing and computer vision; however, the SoftMax operation has become a major bottleneck in Transformer inference due to costly exponential operations and repeated memory access. We present a hardware-native SoftMax engine based on an asynchronous probabilistic Ising machine built from experimentally fabricated SOT-MTJ P-Bit. A SoftMax-To-Ising compiler maps arbitrary logits into an implementable energy function, enabling direct physical sampling of the corresponding Boltzmann distribution. The resulting Ising SoftMax converges rapidly, exceeding FP16 precision at 32μ text {s}. When integrated into neural network workloads, it preserves high inference accuracy without noticeable degradation. These results demonstrate a fully asynchronous physical SoftMax engine and highlight its potential for energy-efficient deep learning acceleration.

Original languageEnglish
Pages (from-to)1458-1461
Number of pages4
JournalIEEE Electron Device Letters
Volume47
Issue number7
DOIs
StatePublished - 1 Jul 2026

Keywords

  • Ising machine
  • P-Bit
  • SOT-MTJ
  • transformer

Fingerprint

Dive into the research topics of 'Energy-Efficient Parallel Ising SoftMax Engine for Transformer Attention Based on Asynchronous SOT-P-Bits'. Together they form a unique fingerprint.

Cite this