TY - JOUR
T1 - Cooperative decision-making algorithm with efficient convergence for UCAV formation in beyond-visual-range air combat based on multi-agent reinforcement learning
T2 - Cooperative decision-making algorithm with efficient convergence
AU - ZHOU, Yaoming
AU - YANG, Fan
AU - ZHANG, Chaoyue
AU - LI, Shida
AU - WANG, Yongchao
N1 - Publisher Copyright:
© 2024
PY - 2024/8
Y1 - 2024/8
N2 - Highly intelligent Unmanned Combat Aerial Vehicle (UCAV) formation is expected to bring out strengths in Beyond-Visual-Range (BVR) air combat. Although Multi-Agent Reinforcement Learning (MARL) shows outstanding performance in cooperative decision-making, it is challenging for existing MARL algorithms to quickly converge to an optimal strategy for UCAV formation in BVR air combat where confrontation is complicated and reward is extremely sparse and delayed. Aiming to solve this problem, this paper proposes an Advantage Highlight Multi-Agent Proximal Policy Optimization (AHMAPPO) algorithm. First, at every step, the AHMAPPO records the degree to which the best formation exceeds the average of formations in parallel environments and carries out additional advantage sampling according to it. Then, the sampling result is introduced into the updating process of the actor network to improve its optimization efficiency. Finally, the simulation results reveal that compared with some state-of-the-art MARL algorithms, the AHMAPPO can obtain a more excellent strategy utilizing fewer sample episodes in the UCAV formation BVR air combat simulation environment built in this paper, which can reflect the critical features of BVR air combat. The AHMAPPO can significantly increase the convergence efficiency of the strategy for UCAV formation in BVR air combat, with a maximum increase of 81.5% relative to other algorithms.
AB - Highly intelligent Unmanned Combat Aerial Vehicle (UCAV) formation is expected to bring out strengths in Beyond-Visual-Range (BVR) air combat. Although Multi-Agent Reinforcement Learning (MARL) shows outstanding performance in cooperative decision-making, it is challenging for existing MARL algorithms to quickly converge to an optimal strategy for UCAV formation in BVR air combat where confrontation is complicated and reward is extremely sparse and delayed. Aiming to solve this problem, this paper proposes an Advantage Highlight Multi-Agent Proximal Policy Optimization (AHMAPPO) algorithm. First, at every step, the AHMAPPO records the degree to which the best formation exceeds the average of formations in parallel environments and carries out additional advantage sampling according to it. Then, the sampling result is introduced into the updating process of the actor network to improve its optimization efficiency. Finally, the simulation results reveal that compared with some state-of-the-art MARL algorithms, the AHMAPPO can obtain a more excellent strategy utilizing fewer sample episodes in the UCAV formation BVR air combat simulation environment built in this paper, which can reflect the critical features of BVR air combat. The AHMAPPO can significantly increase the convergence efficiency of the strategy for UCAV formation in BVR air combat, with a maximum increase of 81.5% relative to other algorithms.
KW - Advantage highlight
KW - Beyond-visual-range (BVR) air combat
KW - Decision-making
KW - Multi-agent reinforcement learning (MARL)
KW - Unmanned combat aerial vehicle (UCAV) formation
UR - https://www.scopus.com/pages/publications/85198384502
U2 - 10.1016/j.cja.2024.04.008
DO - 10.1016/j.cja.2024.04.008
M3 - 文章
AN - SCOPUS:85198384502
SN - 1000-9361
VL - 37
SP - 311
EP - 328
JO - Chinese Journal of Aeronautics
JF - Chinese Journal of Aeronautics
IS - 8
ER -