跳到主要导航 跳到搜索 跳到主要内容

Reinforcement learning solution for HJB equation arising in constrained optimal control problem

  • Biao Luo
  • , Huai Ning Wu
  • , Tingwen Huang
  • , Derong Liu*
  • *此作品的通讯作者
  • Chinese Academy of Sciences
  • Texas A&M University at Qatar
  • University of Science and Technology Beijing

科研成果: 期刊稿件文章同行评审

摘要

The constrained optimal control problem depends on the solution of the complicated Hamilton-Jacobi-Bellman equation (HJBE). In this paper, a data-based off-policy reinforcement learning (RL) method is proposed, which learns the solution of the HJBE and the optimal control policy from real system data. One important feature of the off-policy RL is that its policy evaluation can be realized with data generated by other behavior policies, not necessarily the target policy, which solves the insufficient exploration problem. The convergence of the off-policy RL is proved by demonstrating its equivalence to the successive approximation approach. Its implementation procedure is based on the actor-critic neural networks structure, where the function approximation is conducted with linearly independent basis functions. Subsequently, the convergence of the implementation procedure with function approximation is also proved. Finally, its effectiveness is verified through computer simulations.

源语言英语
页(从-至)150-158
页数9
期刊Neural Networks
71
DOI
出版状态已出版 - 1 11月 2015

指纹

探究 'Reinforcement learning solution for HJB equation arising in constrained optimal control problem' 的科研主题。它们共同构成独一无二的指纹。

引用此