TY - GEN
T1 - Data Driven Evasion Policy for Repeated Pursuit Evasion Problem with Unknown Imperfect Pursuer
AU - He, Linkun
AU - Zhang, Ran
AU - Li, Huifeng
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - A novel repeated pursuit-evasion (PE) problem with unknown imperfect pursuer is proposed and addressed using model based reinforcement learning. Grounded in aerospace applications, the problem adopts a repeated game formulation, where the evader engages repeatedly with the pursuer with unknown imperfect policy for multiple rounds, and the objective for the evader is to optimize the evasion policy based on data of preceding rounds so that evasion is achieved with minimum number of rounds. Unlike previous works that rely on prior knowledge of the pursuer's policy, our approach focus on efficient data exploitation in each round. Specifically, we adopt the probabilistic inference for learning control (PILCO) method, where a Gaussian process is adopted to model the unknown imperfect pursuer. This allows for efficient data exploitation and transforms the problem into the deterministic pay-off optimization with analytical gradients, enabling successive policy optimization in each round. Numerical results on a typical aerospace application demonstrate that the proposed method achieves successful evasion in remarkably few rounds, with the optimized policy aligning closely with the analytical optimal solution assuming complete knowledge of the pursuer.
AB - A novel repeated pursuit-evasion (PE) problem with unknown imperfect pursuer is proposed and addressed using model based reinforcement learning. Grounded in aerospace applications, the problem adopts a repeated game formulation, where the evader engages repeatedly with the pursuer with unknown imperfect policy for multiple rounds, and the objective for the evader is to optimize the evasion policy based on data of preceding rounds so that evasion is achieved with minimum number of rounds. Unlike previous works that rely on prior knowledge of the pursuer's policy, our approach focus on efficient data exploitation in each round. Specifically, we adopt the probabilistic inference for learning control (PILCO) method, where a Gaussian process is adopted to model the unknown imperfect pursuer. This allows for efficient data exploitation and transforms the problem into the deterministic pay-off optimization with analytical gradients, enabling successive policy optimization in each round. Numerical results on a typical aerospace application demonstrate that the proposed method achieves successful evasion in remarkably few rounds, with the optimized policy aligning closely with the analytical optimal solution assuming complete knowledge of the pursuer.
KW - Gaussian process
KW - pursuit evasion game
KW - reinforcement learning
UR - https://www.scopus.com/pages/publications/105036001758
U2 - 10.1109/RAAI67517.2025.11423324
DO - 10.1109/RAAI67517.2025.11423324
M3 - 会议稿件
AN - SCOPUS:105036001758
T3 - 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence, RAAI 2025
SP - 6
EP - 10
BT - 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence, RAAI 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence, RAAI 2025
Y2 - 18 December 2025 through 20 December 2025
ER -