跳到主要导航 跳到搜索 跳到主要内容

A scalable reinforcement learning framework for multi-UAV autonomous soaring in stochastic updraft environments

  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Inspired by soaring birds that harvest energy from thermal updrafts for sustained flight, unpowered UAVs have increasingly adopted thermal exploitation strategies to extend endurance, and recent studies have further expanded these approaches toward multi-UAV operation in stochastic environments. However, existing reinforcement learning methods for multi-UAV thermal exploitation often suffer from limited scalability, strong reliance on explicit inter-agent interaction structures, and physically inconsistent high-frequency control actions, which restrict their effectiveness in realistic scenarios. This paper investigates a bio-inspired thermal searching and soaring problem and proposes Repeated Action and Node Encoding Proximal Policy Optimization (RANE-PPO), an improved multi-agent reinforcement learning framework. A realistic simulation environment integrates a Gaussian-based stochastic thermal model with a frigatebird-inspired UAV dynamic model to capture energy-harvesting characteristics. RANE-PPO employs a node-encoding architecture that independently encodes each agent’s local observations, augmented with auxiliary updraft-related information accumulated during exploration, without relying on explicit global coordinate information of all agents or predefined inter-agent structural relationships. In addition, an action repetition mechanism is incorporated to enforce temporal consistency in control actions. Simulation results show that RANE-PPO achieves faster convergence, higher cumulative rewards, and improved long-duration flight success rates compared with PPO and Graphic Neural Network based PPO, while maintaining strong robustness and generalization under varying and even increased numbers of agents. These results indicate that RANE-PPO provides an effective and scalable solution for multi agent thermal searching and soaring in complex and uncertain environments.

源语言英语
文章编号112488
期刊Aerospace Science and Technology
178
DOI
出版状态已出版 - 11月 2026

指纹

探究 'A scalable reinforcement learning framework for multi-UAV autonomous soaring in stochastic updraft environments' 的科研主题。它们共同构成独一无二的指纹。

引用此