跳到主要导航 跳到搜索 跳到主要内容

High-Dimensional Hyperparameter Optimization via Adjoint Differentiation

  • Beihang University
  • Hunan University
  • Peng Cheng Laboratory
  • Zhongguancun Academy
  • Academy of Military Medical Science China

科研成果: 期刊稿件文章同行评审

摘要

As an emerging machine learning task, high-dimensional hyperparameter optimization (HO) aims at enhancing traditional deep learning models by simultaneously optimizing the neural networks' weights and hyperparameters in a joint bilevel configuration. However, such nested objectives can impose nontrivial difficulties for the pursuit of the gradient of the validation risk with respect to the hyperparameters (a.k.a. hypergradient). To tackle this challenge, we revisit its bilevel objective from the novel perspective of continuous dynamics and then solve the whole HO problem with the adjoint state theory. The proposed HO framework, termed Adjoint Diff, is naturally scalable to a very deep neural network with high-dimensional hyperparameters because it only requires constant memory cost in training. Adjoint Diff is in fact, a general framework that some existing gradient-based HO algorithms are well interpreted by it with simple algebra. In addition, we further offer the Adjoint Diff+ framework by incorporating the prevalent momentum learning concept into the basic Adjoint Diff for enhanced convergence. Experimental results show that our Adjoint Diff frameworks outperform several state-of-the-art approaches on three high-dimensional HO instances including, designing a loss function for imbalanced data, selecting samples from noisy labels, and learning auxiliary tasks for fine-grained classification.

源语言英语
页(从-至)2148-2162
页数15
期刊IEEE Transactions on Artificial Intelligence
6
8
DOI
出版状态已出版 - 2025

学术指纹

探究 'High-Dimensional Hyperparameter Optimization via Adjoint Differentiation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此