TY - GEN
T1 - A Scale-Independent Deep Reinforcement Learning Framework for Multi-UAV Communication Resource Allocation in Unmanned Logistics
AU - Huang, Lizhen
AU - Wang, Feipeng
AU - Liu, Chunhui
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Unmanned aerial vehicles (UAVs) have emerged as a promising solution for enhancing the effectiveness and accuracy of logistics distribution in inconvenient scenarios, such as disaster areas and remote regions. Wireless resource (spectrum, power) allocation is crucial for the multi-UAV system to enhance the quality of the service. However, obtaining an optimal strategy for multi-UAV resource allocation is challenging, especially when the number of UAVs varies. To address this problem, a scale-independent deep reinforcement learning (DRL) framework is proposed in this paper, which can adapt to different DRL models. Firstly, a decentralized Partially Observable Markov Decision Process (dec-POMDP) model is presented for multi-UAV communication networks, in which each UAV's state involves partial observation and local communication. Secondly, to improve the scalability, a state representation scheme is proposed to integrate the variable-dimension information set into a fixed-shape input variable while maintaining the permutation irrelevance. Finally, three DRL methods are incorporated with the state representation scheme to resolve the dec-POMDP problem. Simulation results demonstrate that the proposed scale-independent DRL framework can learn a cooperative resource allocation policy. Furthermore, the DRL-based methods outperform the greedy-based and random methods in terms of communication capacity.
AB - Unmanned aerial vehicles (UAVs) have emerged as a promising solution for enhancing the effectiveness and accuracy of logistics distribution in inconvenient scenarios, such as disaster areas and remote regions. Wireless resource (spectrum, power) allocation is crucial for the multi-UAV system to enhance the quality of the service. However, obtaining an optimal strategy for multi-UAV resource allocation is challenging, especially when the number of UAVs varies. To address this problem, a scale-independent deep reinforcement learning (DRL) framework is proposed in this paper, which can adapt to different DRL models. Firstly, a decentralized Partially Observable Markov Decision Process (dec-POMDP) model is presented for multi-UAV communication networks, in which each UAV's state involves partial observation and local communication. Secondly, to improve the scalability, a state representation scheme is proposed to integrate the variable-dimension information set into a fixed-shape input variable while maintaining the permutation irrelevance. Finally, three DRL methods are incorporated with the state representation scheme to resolve the dec-POMDP problem. Simulation results demonstrate that the proposed scale-independent DRL framework can learn a cooperative resource allocation policy. Furthermore, the DRL-based methods outperform the greedy-based and random methods in terms of communication capacity.
KW - Deep reinforcement learning
KW - Multi-UAV communication networks
KW - Resource allocation
KW - Scale-Independent
UR - https://www.scopus.com/pages/publications/85186077273
U2 - 10.1109/ITAIC58329.2023.10409035
DO - 10.1109/ITAIC58329.2023.10409035
M3 - 会议稿件
AN - SCOPUS:85186077273
T3 - IEEE Joint International Information Technology and Artificial Intelligence Conference (ITAIC)
SP - 1609
EP - 1618
BT - IEEE ITAIC 2023 - IEEE 11th Joint International Information Technology and Artificial Intelligence Conference
A2 - Xu, Bing
A2 - Mou, Kefen
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 11th Joint International Information Technology and Artificial Intelligence Conference, ITAIC 2023
Y2 - 8 December 2023 through 10 December 2023
ER -