TY - GEN
T1 - Dynamic Weather Avoidance Based on Maximum Diffusion Reinforcement Learning
AU - Wang, Yuxin
AU - Cai, Kaiquan
AU - Zhao, Peng
AU - Xu, Hanjie
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - This paper investigates an autonomous trajectory replanning framework for aircraft operations based on Maximum Diffusion Reinforcement Learning. The method fully considers the complex coupling between aircraft high-inertia kinematics and hazardous weather obstacles, addressing the critical exploration failure challenge inherent in conventional Deep Reinforcement Learning (DRL) algorithms. Specifically, high-fidelity flight dynamics based on the Base of Aircraft Data (BADA) model are integrated into the decision-making process. To overcome the exploration limitations of standard Maximum Entropy RL caused by strong temporal correlations in state transitions, a trajectory entropy maximization mechanism is introduced. This mechanism utilizes the determinant of the trajectory autocovariance matrix as a regularization term to enforce diverse state-space coverage. Numerical results and ablation studies demonstrate that the proposed framework significantly outperforms the baseline Soft Actor-Critic (SAC) algorithm. In scenarios with complex obstacle distributions, the proposed method achieves and higher collision avoidance success rates. Furthermore, the ablation analysis verifies the critical role of the diffusion regularization term in preventing policy collapse and ensuring robust dynamic replanning capabilities.
AB - This paper investigates an autonomous trajectory replanning framework for aircraft operations based on Maximum Diffusion Reinforcement Learning. The method fully considers the complex coupling between aircraft high-inertia kinematics and hazardous weather obstacles, addressing the critical exploration failure challenge inherent in conventional Deep Reinforcement Learning (DRL) algorithms. Specifically, high-fidelity flight dynamics based on the Base of Aircraft Data (BADA) model are integrated into the decision-making process. To overcome the exploration limitations of standard Maximum Entropy RL caused by strong temporal correlations in state transitions, a trajectory entropy maximization mechanism is introduced. This mechanism utilizes the determinant of the trajectory autocovariance matrix as a regularization term to enforce diverse state-space coverage. Numerical results and ablation studies demonstrate that the proposed framework significantly outperforms the baseline Soft Actor-Critic (SAC) algorithm. In scenarios with complex obstacle distributions, the proposed method achieves and higher collision avoidance success rates. Furthermore, the ablation analysis verifies the critical role of the diffusion regularization term in preventing policy collapse and ensuring robust dynamic replanning capabilities.
KW - Autonomous operations
KW - Base of Aircraft Data (BADA)
KW - dynamic weather avoidance
KW - Maximum Diffusion Reinforcement Learning
UR - https://www.scopus.com/pages/publications/105043618804
U2 - 10.1109/ICNS69853.2026.11570271
DO - 10.1109/ICNS69853.2026.11570271
M3 - 会议稿件
AN - SCOPUS:105043618804
T3 - Integrated Communications, Navigation and Surveillance Conference, ICNS
BT - ICNS 2026 - 2026 Integrated Communications, Navigation and Surveillance Conference, Conference Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 26th Integrated Communications, Navigation and Surveillance, ICNS 2026
Y2 - 14 April 2026 through 16 April 2026
ER -