TY - JOUR
T1 - A scalable reinforcement learning framework for multi-UAV autonomous soaring in stochastic updraft environments
AU - Qiu, Wenwei
AU - Wan, Zhiqiang
AU - Yan, De
AU - Dong, Ruihan
N1 - Publisher Copyright:
© 2026 The Authors.
PY - 2026/11
Y1 - 2026/11
N2 - Inspired by soaring birds that harvest energy from thermal updrafts for sustained flight, unpowered UAVs have increasingly adopted thermal exploitation strategies to extend endurance, and recent studies have further expanded these approaches toward multi-UAV operation in stochastic environments. However, existing reinforcement learning methods for multi-UAV thermal exploitation often suffer from limited scalability, strong reliance on explicit inter-agent interaction structures, and physically inconsistent high-frequency control actions, which restrict their effectiveness in realistic scenarios. This paper investigates a bio-inspired thermal searching and soaring problem and proposes Repeated Action and Node Encoding Proximal Policy Optimization (RANE-PPO), an improved multi-agent reinforcement learning framework. A realistic simulation environment integrates a Gaussian-based stochastic thermal model with a frigatebird-inspired UAV dynamic model to capture energy-harvesting characteristics. RANE-PPO employs a node-encoding architecture that independently encodes each agent’s local observations, augmented with auxiliary updraft-related information accumulated during exploration, without relying on explicit global coordinate information of all agents or predefined inter-agent structural relationships. In addition, an action repetition mechanism is incorporated to enforce temporal consistency in control actions. Simulation results show that RANE-PPO achieves faster convergence, higher cumulative rewards, and improved long-duration flight success rates compared with PPO and Graphic Neural Network based PPO, while maintaining strong robustness and generalization under varying and even increased numbers of agents. These results indicate that RANE-PPO provides an effective and scalable solution for multi agent thermal searching and soaring in complex and uncertain environments.
AB - Inspired by soaring birds that harvest energy from thermal updrafts for sustained flight, unpowered UAVs have increasingly adopted thermal exploitation strategies to extend endurance, and recent studies have further expanded these approaches toward multi-UAV operation in stochastic environments. However, existing reinforcement learning methods for multi-UAV thermal exploitation often suffer from limited scalability, strong reliance on explicit inter-agent interaction structures, and physically inconsistent high-frequency control actions, which restrict their effectiveness in realistic scenarios. This paper investigates a bio-inspired thermal searching and soaring problem and proposes Repeated Action and Node Encoding Proximal Policy Optimization (RANE-PPO), an improved multi-agent reinforcement learning framework. A realistic simulation environment integrates a Gaussian-based stochastic thermal model with a frigatebird-inspired UAV dynamic model to capture energy-harvesting characteristics. RANE-PPO employs a node-encoding architecture that independently encodes each agent’s local observations, augmented with auxiliary updraft-related information accumulated during exploration, without relying on explicit global coordinate information of all agents or predefined inter-agent structural relationships. In addition, an action repetition mechanism is incorporated to enforce temporal consistency in control actions. Simulation results show that RANE-PPO achieves faster convergence, higher cumulative rewards, and improved long-duration flight success rates compared with PPO and Graphic Neural Network based PPO, while maintaining strong robustness and generalization under varying and even increased numbers of agents. These results indicate that RANE-PPO provides an effective and scalable solution for multi agent thermal searching and soaring in complex and uncertain environments.
KW - Bio-inspired UAV
KW - Multi-UAV soaring
KW - Proximal policy optimization
KW - Reinforcement learning
KW - Thermal updraft
UR - https://www.scopus.com/pages/publications/105039087450
U2 - 10.1016/j.ast.2026.112488
DO - 10.1016/j.ast.2026.112488
M3 - 文章
AN - SCOPUS:105039087450
SN - 1270-9638
VL - 178
JO - Aerospace Science and Technology
JF - Aerospace Science and Technology
M1 - 112488
ER -