Abstract
Sharpness-Aware Minimization (SAM) improves model generalization but doubles the computational cost of Stochastic Gradient Descent (SGD) by requiring twice the gradient calculations per optimization step. To mitigate this, we propose Adaptively sampling-Reusing-mixing decomposed gradients to significantly accelerate SAM (ARSAM). Concretely, we discover that SAM's gradient can be decomposed into the SGD gradient and the Projection of the Second-order gradient onto the First-order gradient (PSF). Furthermore, we observe that the PSF plays an increasingly critical role in achieving flatter minima. Therefore, ARSAM is proposed to adaptively adjust the computation frequency of the PSF and reuse historical PSF, ensuring that the model maintains strong generalization ability while significantly reducing computational overhead. Extensive experiments show that ARSAM achieves state-of-the-art accuracies comparable to SAM across diverse network architectures. On CIFAR-10/100, ARSAM is comparable to SAM while providing a speedup of about 40%. Moreover, ARSAM accelerates optimization for the various challenge tasks (e.g., human pose estimation, and model quantization) without sacrificing performance, demonstrating its broad practicality.
| Original language | English |
|---|---|
| Article number | 114031 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| State | Published - Dec 2026 |
Keywords
- Accelerate
- Decompose
- Generalization
- Sharpness-aware
Fingerprint
Dive into the research topics of 'Adaptively sampling-reusing-mixing decomposed gradients to speed up sharpness aware minimization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver