TY - GEN
T1 - Role-Structured Reinforcement Learning for Heterogeneous Airship-Guided UAV Pursuit Tracking
AU - Yang, Kejie
AU - Zhou, Yuting
AU - Zhu, Ming
AU - Zhang, Yifei
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - This paper investigates heterogeneous aerial pursuit tracking with one high-endurance airship, multiple UAV pursuers, and evasive targets. The task is formulated as a partially observed mixed cooperative-adversarial game under centralized training and decentralized execution. A role-difference-potential multi-agent reinforcement learning framework is proposed, in which each agent maintains an individual actor and critics are shared at the role level with identity conditioning. To improve cooperative credit assignment, pursuer optimization combines potential-based shaping with counterfactual one-step difference rewards. The potential function integrates coverage guidance, approach pressure, capture-neighborhood entry, encirclement quality, and safety regularization; actions are parameterized by bounded speed and yaw rate for kinematic consistency. Simulation results show stable coordinated closure and reliable capture, indicating that role-structured entropy-regularized learning with difference-potential rewards is a practical baseline for heterogeneous pursuit tasks.
AB - This paper investigates heterogeneous aerial pursuit tracking with one high-endurance airship, multiple UAV pursuers, and evasive targets. The task is formulated as a partially observed mixed cooperative-adversarial game under centralized training and decentralized execution. A role-difference-potential multi-agent reinforcement learning framework is proposed, in which each agent maintains an individual actor and critics are shared at the role level with identity conditioning. To improve cooperative credit assignment, pursuer optimization combines potential-based shaping with counterfactual one-step difference rewards. The potential function integrates coverage guidance, approach pressure, capture-neighborhood entry, encirclement quality, and safety regularization; actions are parameterized by bounded speed and yaw rate for kinematic consistency. Simulation results show stable coordinated closure and reliable capture, indicating that role-structured entropy-regularized learning with difference-potential rewards is a practical baseline for heterogeneous pursuit tasks.
KW - Airship-UAV Cooperation
KW - Heterogeneous Multi-Agent Reinforcement Learning
KW - PotentialBased Reward Shaping
KW - Pursuit-Evasion
UR - https://www.scopus.com/pages/publications/105041682904
U2 - 10.1109/RPIC68328.2026.11518352
DO - 10.1109/RPIC68328.2026.11518352
M3 - 会议稿件
AN - SCOPUS:105041682904
T3 - 2026 International Conference on Robot Perception and Intelligent Control, RPIC 2026
SP - 96
EP - 100
BT - 2026 International Conference on Robot Perception and Intelligent Control, RPIC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 International Conference on Robot Perception and Intelligent Control, RPIC 2026
Y2 - 27 March 2026 through 29 March 2026
ER -