TY - JOUR
T1 - A Spatiotemporal Graph Reasoning Approach for Pursuit-Evasion Game With Communication Limits
AU - Yang, Zhengzhi
AU - Du, Wenbo
AU - Zhang, Xin
AU - Wang, Jingjing
AU - Li, Yumeng
N1 - Publisher Copyright:
© 1967-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Uncrewed aerial vehicle (UAV) swarms performing encirclement tasks require reliable communications to coordinate strategies to capture the fleeter target. However, dense urban buildings can obstruct signals, causing communication disruptions, observation fluctuations, and decreased strategy efficiency. To address this problem, we propose the Relational Graph Attention Twin Delayed DDPG (RGATD3) approach. In particular, we adopt a spatiotemporal graph to represent complex state information, and a graph prediction network to anticipate missing communication data, effectively stabilizing observation fluctuations. In order to improve decision-making in complex environments, a relational graph learning mechanism is introduced, where Graph Attention Networks (GATs) extract agent features based on specific relational types, and a multi-head self-attention mechanism aggregates these features. Combined with localized detection, this approach enhances decision-making from complex observation. Additionally, well-tailored reward functions further encourage agents to restore communication and cooperation. The results indicate that RGATD3 outperforms existing state-of-the-art methods, increasing encirclement success and reducing capture time under fluctuating communication conditions.
AB - Uncrewed aerial vehicle (UAV) swarms performing encirclement tasks require reliable communications to coordinate strategies to capture the fleeter target. However, dense urban buildings can obstruct signals, causing communication disruptions, observation fluctuations, and decreased strategy efficiency. To address this problem, we propose the Relational Graph Attention Twin Delayed DDPG (RGATD3) approach. In particular, we adopt a spatiotemporal graph to represent complex state information, and a graph prediction network to anticipate missing communication data, effectively stabilizing observation fluctuations. In order to improve decision-making in complex environments, a relational graph learning mechanism is introduced, where Graph Attention Networks (GATs) extract agent features based on specific relational types, and a multi-head self-attention mechanism aggregates these features. Combined with localized detection, this approach enhances decision-making from complex observation. Additionally, well-tailored reward functions further encourage agents to restore communication and cooperation. The results indicate that RGATD3 outperforms existing state-of-the-art methods, increasing encirclement success and reducing capture time under fluctuating communication conditions.
KW - Pursuit-evasion game (PEG)
KW - cooperative encirclement
KW - graph attention networks (GATs)
KW - multi-agent reinforcement learning (MARL)
KW - relational graph learning (RGL)
UR - https://www.scopus.com/pages/publications/105014981744
U2 - 10.1109/TVT.2025.3605736
DO - 10.1109/TVT.2025.3605736
M3 - 文章
AN - SCOPUS:105014981744
SN - 0018-9545
VL - 75
SP - 2069
EP - 2085
JO - IEEE Transactions on Vehicular Technology
JF - IEEE Transactions on Vehicular Technology
IS - 2
ER -