TY - GEN
T1 - Battlefield Situation Deduction and Maneuver Decision Using Deep Q-Learning
AU - Shi, Minghui
AU - Dong, Xiwang
AU - Han, Liang
AU - Li, Qingdong
AU - Ren, Zhang
N1 - Publisher Copyright:
© 2021 Technical Committee on Control Theory, Chinese Association of Automation.
PY - 2021/7/26
Y1 - 2021/7/26
N2 - As the pace and complexity of modern warfare accelerates, it is of great significance to apply intelligent technology to defense decision-making. Due to the high dynamics and randomness of the aircraft, traditional methods are difficult to solve the optimal control strategy. The characteristics of reinforcement learning match the difficulty of the problem. In this paper, the situation of hypersonic aircraft is deduced, and the deep reinforcement learning method is used to make autonomous penetration decisions in the reentry phase. The model of aircraft and environment is established, and the maneuvering decision-making model is established based on deep Q-learning and its optimization algorithm. Through a large number of simulation training, this method can effectively give the real-time decision output of the agent and make a good prediction of the situation. It has the ability of short-range accurate operation and long-term planning and prediction. This method can improve the probability of successful penetration, and can be used as the decision-making basis of glider penetration.
AB - As the pace and complexity of modern warfare accelerates, it is of great significance to apply intelligent technology to defense decision-making. Due to the high dynamics and randomness of the aircraft, traditional methods are difficult to solve the optimal control strategy. The characteristics of reinforcement learning match the difficulty of the problem. In this paper, the situation of hypersonic aircraft is deduced, and the deep reinforcement learning method is used to make autonomous penetration decisions in the reentry phase. The model of aircraft and environment is established, and the maneuvering decision-making model is established based on deep Q-learning and its optimization algorithm. Through a large number of simulation training, this method can effectively give the real-time decision output of the agent and make a good prediction of the situation. It has the ability of short-range accurate operation and long-term planning and prediction. This method can improve the probability of successful penetration, and can be used as the decision-making basis of glider penetration.
KW - deep Q-learning
KW - Hypersonic aircraft
KW - maneuver decision-making
KW - reinforcement learning
KW - situational deduction
UR - https://www.scopus.com/pages/publications/85117321573
U2 - 10.23919/CCC52363.2021.9549649
DO - 10.23919/CCC52363.2021.9549649
M3 - 会议稿件
AN - SCOPUS:85117321573
T3 - Chinese Control Conference, CCC
SP - 3651
EP - 3656
BT - Proceedings of the 40th Chinese Control Conference, CCC 2021
A2 - Peng, Chen
A2 - Sun, Jian
PB - IEEE Computer Society
T2 - 40th Chinese Control Conference, CCC 2021
Y2 - 26 July 2021 through 28 July 2021
ER -