摘要
Sharpness-Aware Minimization (SAM) improves model generalization but doubles the computational cost of Stochastic Gradient Descent (SGD) by requiring twice the gradient calculations per optimization step. To mitigate this, we propose Adaptively sampling-Reusing-mixing decomposed gradients to significantly accelerate SAM (ARSAM). Concretely, we discover that SAM's gradient can be decomposed into the SGD gradient and the Projection of the Second-order gradient onto the First-order gradient (PSF). Furthermore, we observe that the PSF plays an increasingly critical role in achieving flatter minima. Therefore, ARSAM is proposed to adaptively adjust the computation frequency of the PSF and reuse historical PSF, ensuring that the model maintains strong generalization ability while significantly reducing computational overhead. Extensive experiments show that ARSAM achieves state-of-the-art accuracies comparable to SAM across diverse network architectures. On CIFAR-10/100, ARSAM is comparable to SAM while providing a speedup of about 40%. Moreover, ARSAM accelerates optimization for the various challenge tasks (e.g., human pose estimation, and model quantization) without sacrificing performance, demonstrating its broad practicality.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 114031 |
| 期刊 | Pattern Recognition |
| 卷 | 180 |
| DOI | |
| 出版状态 | 已出版 - 12月 2026 |
学术指纹
探究 'Adaptively sampling-reusing-mixing decomposed gradients to speed up sharpness aware minimization' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver