跳到主要导航 跳到搜索 跳到主要内容

用于深度学习训练加速的自适应框架设计

  • Beihang University
  • DeePoly Technology Inc.

科研成果: 期刊稿件文章同行评审

摘要

Field-programmable gate array (FPGA) is usually used to accelerate the training phase of deep learning algo-rithms, but it usually requires a long development cycle and rich hardware design expertise for satisfied execution performance. In order to deal with this challenge, an adaptive acceleration framework for deep learning algorithm is proposed in this paper. We investigate the application scale, parallel scheduling strategy, resource usage and the scalability of functionality. With the CPU-FPGA heterogeneous acceleration template based technology, an adaptive model compiler is proposed to customize the accelerator based on the algorithm's complexity and hardware resources available. The proposed hardware and software co-design framework can effectively adapt to different FPGA hardware resources and support the fast evolution of deep learning algorithms. Taking the graph neural network as an example, it can obtain 7~41x performance improvements compared to the general purpose CPU platform.

投稿的翻译标题Template-Based Adaptive Training Acceleration Framework for Deep Learning Algorithms
源语言繁体中文
页(从-至)974-982
页数9
期刊Jisuanji Fuzhu Sheji Yu Tuxingxue Xuebao/Journal of Computer-Aided Design and Computer Graphics
33
6
DOI
出版状态已出版 - 20 6月 2021

关键词

  • Deep learning
  • Field-programmable
  • Graph convolutional networks (GCN)
  • Heterogeneous accelerator

学术指纹

探究 '用于深度学习训练加速的自适应框架设计' 的科研主题。它们共同构成独一无二的学术指纹。

引用此