TY - JOUR
T1 - Iterative adaptive dynamic programming methods with neural network implementation for multi-player zero-sum games
AU - Jiang, He
AU - Zhang, Huaguang
AU - Han, Ji
AU - Zhang, Kun
N1 - Publisher Copyright:
© 2018
PY - 2018/9/13
Y1 - 2018/9/13
N2 - This paper presents novel iterative learning methods along with the neural network implementation for multi-player zero-sum games. Solving zero-sum games depends on the solutions of Hamilton–Jacobi-Isaacs equations, which are nonlinear partial differential equations. These solutions are generally difficult or even impossible to be obtained analytically. To overcome this difficulty, iterative adaptive dynamic programming algorithms are utilized. In the related research works, three-network architecture, i.e., critic-actor-disturbance structure, is used to approximate the value function, control policies and disturbance policies. Different from the previous works, this paper employs single-network architecture, i.e., critic-only structure, to implement the proposed algorithms, which reduces the computation burden and the complexity of design procedure. Finally, two simulation examples are provided to illustrate the effectiveness of our proposed methods.
AB - This paper presents novel iterative learning methods along with the neural network implementation for multi-player zero-sum games. Solving zero-sum games depends on the solutions of Hamilton–Jacobi-Isaacs equations, which are nonlinear partial differential equations. These solutions are generally difficult or even impossible to be obtained analytically. To overcome this difficulty, iterative adaptive dynamic programming algorithms are utilized. In the related research works, three-network architecture, i.e., critic-actor-disturbance structure, is used to approximate the value function, control policies and disturbance policies. Different from the previous works, this paper employs single-network architecture, i.e., critic-only structure, to implement the proposed algorithms, which reduces the computation burden and the complexity of design procedure. Finally, two simulation examples are provided to illustrate the effectiveness of our proposed methods.
KW - Adaptive dynamic programming
KW - Approximate dynamic programming
KW - Neural networks
KW - Zero-sum games
UR - https://www.scopus.com/pages/publications/85047267198
U2 - 10.1016/j.neucom.2018.04.005
DO - 10.1016/j.neucom.2018.04.005
M3 - 文章
AN - SCOPUS:85047267198
SN - 0925-2312
VL - 307
SP - 54
EP - 60
JO - Neurocomputing
JF - Neurocomputing
ER -