TY - GEN
T1 - A UAV Path Planning Method in Three-Dimensional Urban Airspace based on Safe Reinforcement Learning
AU - Li, Yan
AU - Zhang, Xuejun
AU - Zhu, Yuanjun
AU - Gao, Ziang
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Under the demand of urban terminal "Last Mile Delivery"scenario, finding a safe and efficient UAV path planning method is a crucial issue of current research. Nowadays, reinforcement learning is widely used in UAV path planning, but it is difficult to ensure the safety of the learning or execution phases due to the lack of hard constraints. Aiming at the constraints above, this paper studies how to combine safety properties with RL algorithm to find a safe path and proposes a safe reinforcement learning method called Shield-DDPG for UAV path planning. In the method, a protection mechanism Shield is mainly introduced to prevent the algorithm from outputting unsafe actions. Further, the state space, action space, and reward function are specifically improved for efficiency and safety. Then we compare the Shield-DDPG algorithm with the DDPG and RRT algorithm in some different scenarios, and the results show that the proposed algorithm has a better performance. With the proposed path planning method, UAV can learn well to efficiently and safely reach the destination via calling the trained policy. This research is of great importance to UAV operations and practical applications in complex urban airspace.
AB - Under the demand of urban terminal "Last Mile Delivery"scenario, finding a safe and efficient UAV path planning method is a crucial issue of current research. Nowadays, reinforcement learning is widely used in UAV path planning, but it is difficult to ensure the safety of the learning or execution phases due to the lack of hard constraints. Aiming at the constraints above, this paper studies how to combine safety properties with RL algorithm to find a safe path and proposes a safe reinforcement learning method called Shield-DDPG for UAV path planning. In the method, a protection mechanism Shield is mainly introduced to prevent the algorithm from outputting unsafe actions. Further, the state space, action space, and reward function are specifically improved for efficiency and safety. Then we compare the Shield-DDPG algorithm with the DDPG and RRT algorithm in some different scenarios, and the results show that the proposed algorithm has a better performance. With the proposed path planning method, UAV can learn well to efficiently and safely reach the destination via calling the trained policy. This research is of great importance to UAV operations and practical applications in complex urban airspace.
KW - Deep Deterministic Policy Gradient (DDPG)
KW - Path Planning
KW - Safe Reinforcement Learning (SRL)
KW - Shield
KW - Unmanned Aerial Vehicle (UAV)
UR - https://www.scopus.com/pages/publications/85178654219
U2 - 10.1109/DASC58513.2023.10311219
DO - 10.1109/DASC58513.2023.10311219
M3 - 会议稿件
AN - SCOPUS:85178654219
T3 - AIAA/IEEE Digital Avionics Systems Conference - Proceedings
BT - DASC 2023 - Digital Avionics Systems Conference, Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 42nd IEEE/AIAA Digital Avionics Systems Conference, DASC 2023
Y2 - 1 October 2023 through 5 October 2023
ER -