摘要
Conv-Transformer Neural Networks (CTNNs) achieve remarkable performance in computer vision tasks but face significant inference bottlenecks on edge devices due to the computational complexity of nonlinear operators and inefficient scheduling of heterogeneous processing elements (PEs). To address these challenges, we propose HCTA, a heterogeneous accelerator with three key innovations. First, a nonlinear acceleration engine using static lookup tables (LUTs) results in a precision loss of less than 0.1%. It achieves 1.44×/12.52×/18.56× speedup for Softmax/LayerNorm/GELU with state-of-the-art methods. Second, a Hybrid Control-Data Flow Architecture improves PE utilization by 1.6× compared to prior methods. Third, a Dataflow-Driven Tile Scheduling (DTS) algorithm enables pipelined execution across heterogeneous engines, achieving 2.46× faster performance than existing scheduling methods. Implemented on the Xilinx ZCU102 FPGA, HCTA delivers a throughput of 1388.19 GOP/s, 1.78× higher energy efficiency than state-of-the-art FPGA accelerators, and 47.76×/3.74× better energy efficiency/performance than the NVIDIA V100 GPU.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 133620 |
| 期刊 | Neurocomputing |
| 卷 | 697 |
| DOI | |
| 出版状态 | 已出版 - 7 10月 2026 |
联合国可持续发展目标
此成果有助于实现下列可持续发展目标:
-
可持续发展目标 7 经济适用的清洁能源
指纹
探究 'HCTA: A heterogeneous conv-transformer networks accelerator overcoming nonlinear operator bottlenecks' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver