TY - GEN
T1 - HCTD3
T2 - 5th International Conference on Wireless Communication, Networking and Internet of Things, WCNIoT 2025
AU - Deng, Yu
AU - Xu, Yong
AU - Liu, Chunhui
AU - Yan, Rui
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Resource allocation for multi-UAV cooperative jamming in modern electronic warfare faces significant challenges due to high-dimensional mixed action spaces, complex constraints, and dynamic environments. To address this, this paper introduces a hierarchical reinforcement learning algorithm, HCTD3. The algorithm mitigates complexity by decomposing the task into two sub-problems: Target selection (discrete actions) and power allocation (continuous actions). It employs the Gumbel-Softmax technique for high-level discrete selection and the Twin Delayed DDPG (TD3) algorithm for low-level continuous optimization, enabling efficient learning in the mixed action space. Experimental results demonstrate that, compared to traditional optimization and single-layer reinforcement learning methods, HCTD3 achieves a 37.2% average improvement in jamming effectiveness and converges 45% faster, showcasing its superior performance.
AB - Resource allocation for multi-UAV cooperative jamming in modern electronic warfare faces significant challenges due to high-dimensional mixed action spaces, complex constraints, and dynamic environments. To address this, this paper introduces a hierarchical reinforcement learning algorithm, HCTD3. The algorithm mitigates complexity by decomposing the task into two sub-problems: Target selection (discrete actions) and power allocation (continuous actions). It employs the Gumbel-Softmax technique for high-level discrete selection and the Twin Delayed DDPG (TD3) algorithm for low-level continuous optimization, enabling efficient learning in the mixed action space. Experimental results demonstrate that, compared to traditional optimization and single-layer reinforcement learning methods, HCTD3 achieves a 37.2% average improvement in jamming effectiveness and converges 45% faster, showcasing its superior performance.
KW - Multi-UAV cooperation
KW - electronic jamming
KW - hierarchical reinforcement learning
KW - mixed action space
KW - resource allocation
UR - https://www.scopus.com/pages/publications/105035548258
U2 - 10.1109/WCNIoT67424.2025.11381343
DO - 10.1109/WCNIoT67424.2025.11381343
M3 - 会议稿件
AN - SCOPUS:105035548258
T3 - 2025 5th International Conference on Wireless Communication, Networking and Internet of Things, WCNIoT 2025
SP - 91
EP - 94
BT - 2025 5th International Conference on Wireless Communication, Networking and Internet of Things, WCNIoT 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 5 November 2025 through 7 November 2025
ER -