Skip to main navigation Skip to search Skip to main content

Adaptively sampling-reusing-mixing decomposed gradients to speed up sharpness aware minimization

  • Jiaxin Deng
  • , Junbiao Pang*
  • , Baochang Zhang
  • *Corresponding author for this work
  • Beijing University of Technology
  • Nanchang Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Sharpness-Aware Minimization (SAM) improves model generalization but doubles the computational cost of Stochastic Gradient Descent (SGD) by requiring twice the gradient calculations per optimization step. To mitigate this, we propose Adaptively sampling-Reusing-mixing decomposed gradients to significantly accelerate SAM (ARSAM). Concretely, we discover that SAM's gradient can be decomposed into the SGD gradient and the Projection of the Second-order gradient onto the First-order gradient (PSF). Furthermore, we observe that the PSF plays an increasingly critical role in achieving flatter minima. Therefore, ARSAM is proposed to adaptively adjust the computation frequency of the PSF and reuse historical PSF, ensuring that the model maintains strong generalization ability while significantly reducing computational overhead. Extensive experiments show that ARSAM achieves state-of-the-art accuracies comparable to SAM across diverse network architectures. On CIFAR-10/100, ARSAM is comparable to SAM while providing a speedup of about 40%. Moreover, ARSAM accelerates optimization for the various challenge tasks (e.g., human pose estimation, and model quantization) without sacrificing performance, demonstrating its broad practicality.

Original languageEnglish
Article number114031
JournalPattern Recognition
Volume180
DOIs
StatePublished - Dec 2026

Keywords

  • Accelerate
  • Decompose
  • Generalization
  • Sharpness-aware

Fingerprint

Dive into the research topics of 'Adaptively sampling-reusing-mixing decomposed gradients to speed up sharpness aware minimization'. Together they form a unique fingerprint.

Cite this