Skip to main navigation Skip to search Skip to main content

Exploiting Input Tensor Dynamics in Activation Checkpointing for Efficient Training on GPU

  • Jianjin Liao
  • , Mingzhen Li
  • , Hailong Yang*
  • , Qingxiao Sun
  • , Biao Sun
  • , Jiwei Hao
  • , Tianyu Feng
  • , Fengwei Yu
  • , Shengdong Chen
  • , Ye Tao
  • , Zicheng Zhang
  • , Zhongzhi Luan
  • , Depei Qian
  • *Corresponding author for this work
  • Beihang University
  • SenseTime Group Limited

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Larger deep learning models usually lead to higher model quality, however with an ever-increasing GPU memory footprint. Although several tensor checkpointing techniques have been proposed to enable training under a restricted GPU memory budget, they fail to exploit the input tensor dynamics due to diverse datasets and subsequent data augmentation, and thus leave the training optimization on table. In this paper, we propose Mimose, an input-aware tensor checkpointing planner respecting the memory budget while enabling efficient model training on GPU. Mimose builds a lightweight but accurate prediction model of GPU memory usage online, without pre-analyzing the model. It generates a tensor checkpointing plan based on per-layer memory prediction and applies it to the training process on the fly. Our experiments show that Mimose achieves superior training throughput compared to state-of-the-art checkpointing frameworks under the same GPU memory budgets.

Original languageEnglish
Title of host publicationProceedings - 2023 IEEE International Parallel and Distributed Processing Symposium, IPDPS 2023
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages156-166
Number of pages11
ISBN (Electronic)9798350337662
DOIs
StatePublished - 2023
Event37th IEEE International Parallel and Distributed Processing Symposium, IPDPS 2023 - St. Petersburg, United States
Duration: 15 May 202319 May 2023

Publication series

NameProceedings - 2023 IEEE International Parallel and Distributed Processing Symposium, IPDPS 2023

Conference

Conference37th IEEE International Parallel and Distributed Processing Symposium, IPDPS 2023
Country/TerritoryUnited States
CitySt. Petersburg
Period15/05/2319/05/23

Keywords

  • GPU memory
  • input dynamics
  • model training
  • tensor checkpointing

Fingerprint

Dive into the research topics of 'Exploiting Input Tensor Dynamics in Activation Checkpointing for Efficient Training on GPU'. Together they form a unique fingerprint.

Cite this