TY - JOUR
T1 - An identifier-actor-optimizer policy learning architecture for optimal control of continuous-time nonlinear systems
AU - Cheng, Lin
AU - Wang, Zhen Bo
AU - Jiang, Fang Hua
AU - Li, Jun Feng
N1 - Publisher Copyright:
© 2020, Science China Press and Springer-Verlag GmbH Germany, part of Springer Nature.
PY - 2020/6/1
Y1 - 2020/6/1
N2 - An intelligent solution method is proposed to achieve real-time optimal control for continuous-time nonlinear systems using a novel identifier-actor-optimizer (IAO) policy learning architecture. In this IAO-based policy learning approach, a dynamical identifier is developed to approximate the unknown part of system dynamics using deep neural networks (DNNs). Then, an indirect-method-based optimizer is proposed to generate high-quality optimal actions for system control considering both the constraints and performance index. Furthermore, a DNN-based actor is developed to approximate the obtained optimal actions and return good initial guesses to the optimizer. In this way, the traditional optimal control methods and state-of-the-art DNN techniques are combined in the IAO-based optimal policy learning method. Compared to the reinforcement learning algorithms with actor-critic architectures that suffer hard reward design and low computational efficiency, the IAO-based optimal policy learning algorithm enjoys fewer user-defined parameters, higher learning speeds, and steadier convergence properties in solving complex continuous-time optimal control problems (OCPs). Simulation results of three space flight control missions are given to substantiate the effectiveness of this IAO-based policy learning strategy and to illustrate the performance of the developed DNN-based optimal control method for continuous-time OCPs.
AB - An intelligent solution method is proposed to achieve real-time optimal control for continuous-time nonlinear systems using a novel identifier-actor-optimizer (IAO) policy learning architecture. In this IAO-based policy learning approach, a dynamical identifier is developed to approximate the unknown part of system dynamics using deep neural networks (DNNs). Then, an indirect-method-based optimizer is proposed to generate high-quality optimal actions for system control considering both the constraints and performance index. Furthermore, a DNN-based actor is developed to approximate the obtained optimal actions and return good initial guesses to the optimizer. In this way, the traditional optimal control methods and state-of-the-art DNN techniques are combined in the IAO-based optimal policy learning method. Compared to the reinforcement learning algorithms with actor-critic architectures that suffer hard reward design and low computational efficiency, the IAO-based optimal policy learning algorithm enjoys fewer user-defined parameters, higher learning speeds, and steadier convergence properties in solving complex continuous-time optimal control problems (OCPs). Simulation results of three space flight control missions are given to substantiate the effectiveness of this IAO-based policy learning strategy and to illustrate the performance of the developed DNN-based optimal control method for continuous-time OCPs.
KW - continous-time nonlinear systems
KW - deep neural net-works
KW - identifier-actor-optimizer architecture
KW - intelligent optimal control
UR - https://www.scopus.com/pages/publications/85083107173
U2 - 10.1007/s11433-019-1481-2
DO - 10.1007/s11433-019-1481-2
M3 - 文章
AN - SCOPUS:85083107173
SN - 1674-7348
VL - 63
JO - Science China: Physics, Mechanics and Astronomy
JF - Science China: Physics, Mechanics and Astronomy
IS - 6
M1 - 264511
ER -