跳到主要导航 跳到搜索 跳到主要内容

HCTA: A heterogeneous conv-transformer networks accelerator overcoming nonlinear operator bottlenecks

  • Yongxiang Cao
  • , Hongxu Jiang*
  • , Wei Wang
  • , Yonghua Zhang
  • , Yixiang Zhang
  • , Yanfei Song
  • , Xinyi Wang
  • *此作品的通讯作者
  • Beihang University
  • Tsinghua University

科研成果: 期刊稿件文章同行评审

摘要

Conv-Transformer Neural Networks (CTNNs) achieve remarkable performance in computer vision tasks but face significant inference bottlenecks on edge devices due to the computational complexity of nonlinear operators and inefficient scheduling of heterogeneous processing elements (PEs). To address these challenges, we propose HCTA, a heterogeneous accelerator with three key innovations. First, a nonlinear acceleration engine using static lookup tables (LUTs) results in a precision loss of less than 0.1%. It achieves 1.44×/12.52×/18.56× speedup for Softmax/LayerNorm/GELU with state-of-the-art methods. Second, a Hybrid Control-Data Flow Architecture improves PE utilization by 1.6× compared to prior methods. Third, a Dataflow-Driven Tile Scheduling (DTS) algorithm enables pipelined execution across heterogeneous engines, achieving 2.46× faster performance than existing scheduling methods. Implemented on the Xilinx ZCU102 FPGA, HCTA delivers a throughput of 1388.19 GOP/s, 1.78× higher energy efficiency than state-of-the-art FPGA accelerators, and 47.76×/3.74× better energy efficiency/performance than the NVIDIA V100 GPU.

源语言英语
文章编号133620
期刊Neurocomputing
697
DOI
出版状态已出版 - 7 10月 2026

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 7 - 经济适用的清洁能源
    可持续发展目标 7 经济适用的清洁能源

指纹

探究 'HCTA: A heterogeneous conv-transformer networks accelerator overcoming nonlinear operator bottlenecks' 的科研主题。它们共同构成独一无二的指纹。

引用此