跳到主要导航 跳到搜索 跳到主要内容

Policy Return: A New Method for Reducing the Number of Experimental Trials in Deep Reinforcement Learning

  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Using the same algorithm and hyperparameter configurations, deep reinforcement learning (DRL) will derive drastically different results from multiple experimental trials, and most of these results are unsatisfactory. Because of the instability of the results, researchers have to perform many trials to confirm an algorithm or a set of hyperparameters in DRL. In this article, we present the policy return method, which is a new design for reducing the number of trials when training a DRL model. This method allows the learned policy to return to a previous state when it becomes divergent or stagnant at any stage of training. When returning, a certain percentage of stochastic data is added to the weights of the neural networks to prevent a repeated decline. Extensive experiments on challenging tasks and various target scores demonstrate that the policy return method can bring about a 10% to 40% reduction in the required number of trials compared with that of the corresponding original algorithm, and a 10% to 30% reduction compared with the state-of-the-art algorithms.

源语言英语
文章编号9298771
页(从-至)228099-228107
页数9
期刊IEEE Access
8
DOI
出版状态已出版 - 2020

学术指纹

探究 'Policy Return: A New Method for Reducing the Number of Experimental Trials in Deep Reinforcement Learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此