跳到主要导航 跳到搜索 跳到主要内容

Deep-Reinforcement-Learning-Based Autonomous UAV Navigation with Sparse Rewards

  • Chao Wang
  • , Jian Wang*
  • , Jingjing Wang
  • , Xudong Zhang
  • *此作品的通讯作者
  • Tsinghua University

科研成果: 期刊稿件文章同行评审

摘要

Unmanned aerial vehicles (UAVs) have the potential in delivering Internet-of-Things (IoT) services from a great height, creating an airborne domain of the IoT. In this article, we address the problem of autonomous UAV navigation in large-scale complex environments by formulating it as a Markov decision process with sparse rewards and propose an algorithm named deep reinforcement learning (RL) with nonexpert helpers (LwH). In contrast to prior RL-based methods that put huge efforts into reward shaping, we adopt the sparse reward scheme, i.e., a UAV will be rewarded if and only if it completes navigation tasks. Using the sparse reward scheme ensures that the solution is not biased toward potentially suboptimal directions. However, having no intermediate rewards hinders the agent from efficient learning since informative states are rarely encountered. To handle the challenge, we assume that a prior policy (nonexpert helper) that might be of poor performance is available to the learning agent. The prior policy plays the role of guiding the agent in exploring the state space by reshaping the behavior policy used for environmental interaction. It also assists the agent in achieving goals by setting dynamic learning objectives with increasing difficulty. To evaluate our proposed method, we construct a simulator for UAV navigation in large-scale complex environments and compare our algorithm with several baselines. Experimental results demonstrate that LwH significantly outperforms the state-of-the-art algorithms handling sparse rewards and yields impressive navigation policies comparable to those learned in the environment with dense rewards.

源语言英语
文章编号8993742
页(从-至)6180-6190
页数11
期刊IEEE Internet of Things Journal
7
7
DOI
出版状态已出版 - 7月 2020
已对外发布

学术指纹

探究 'Deep-Reinforcement-Learning-Based Autonomous UAV Navigation with Sparse Rewards' 的科研主题。它们共同构成独一无二的学术指纹。

引用此