跳到主要导航 跳到搜索 跳到主要内容

Balancing Value Iteration and Policy Iteration for Discrete-Time Control

  • Biao Luo*
  • , Yin Yang
  • , Huai Ning Wu
  • , Tingwen Huang
  • *此作品的通讯作者
  • School of Automation
  • Hamad bin Khalifa University
  • Texas A&M University at Qatar

科研成果: 期刊稿件文章同行评审

摘要

The optimal control problem of discrete-time nonlinear systems depends on the solution of the Bellman equation. In this paper, an adaptive reinforcement learning (RL) method is developed to solve the complex Bellman equation, which balances value iteration (VI) and policy iteration (PI). By adding a balance parameter, an adaptive RL integrates VI and PI together, which accelerates VI and avoids the need of an initial admissible control. The convergence of the adaptive RL is proved by showing that it converges to the Bellman equation. Subsequently, the adaptive RL is realized by using the neural network (NN) approximation for value function and a least-squares scheme is developed for updating NN weights. Then, the convergence of NN-based adaptive RL is proved with considering NN approximation error. To further improve its performance, an adaptive rule is developed for tuning balance parameter in adaptive RL iteration by iteration. Finally, the effectiveness of the adaptive RL is validated with simulation studies.

源语言英语
文章编号8657988
页(从-至)3948-3958
页数11
期刊IEEE Transactions on Systems, Man, and Cybernetics: Systems
50
11
DOI
出版状态已出版 - 11月 2020

指纹

探究 'Balancing Value Iteration and Policy Iteration for Discrete-Time Control' 的科研主题。它们共同构成独一无二的指纹。

引用此