跳到主要导航 跳到搜索 跳到主要内容

Self-Imitation Learning for Robot Tasks with Sparse and Delayed Rewards

  • Beihang University
  • Nanjing Research Institute of Electronics Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The application of reinforcement learning (RL) in robotic control is still limited in the environments with sparse and delayed rewards. In this paper, we propose a practical self-imitation learning method named Self-Imitation Learning with Constant Reward (SILCR). Instead of requiring hand-defined immediate rewards from environments, our method assigns the immediate rewards at each timestep with constant values according to their final episodic rewards. In this way, even if the dense rewards from environments are unavailable, every action taken by the agents would be guided properly. We demonstrate the effectiveness of our method in some challenging continuous robotics control tasks in MuJoCo simulation and the results show that our method significantly outperforms the alternative methods in tasks with sparse and delayed rewards. Even compared with alternatives with dense rewards available, our method achieves competitive performance. The ablation experiments also show the stability and reproducibility of our method.

源语言英语
主期刊名2021 IEEE International Conference on Mechatronics and Automation, ICMA 2021
出版商Institute of Electrical and Electronics Engineers Inc.
477-482
页数6
ISBN(电子版)9781665441001
DOI
出版状态已出版 - 8 8月 2021
活动18th IEEE International Conference on Mechatronics and Automation, ICMA 2021 - Takamatsu, 日本
期限: 8 8月 202111 8月 2021

出版系列

姓名2021 IEEE International Conference on Mechatronics and Automation, ICMA 2021

会议

会议18th IEEE International Conference on Mechatronics and Automation, ICMA 2021
国家/地区日本
Takamatsu
时期8/08/2111/08/21

学术指纹

探究 'Self-Imitation Learning for Robot Tasks with Sparse and Delayed Rewards' 的科研主题。它们共同构成独一无二的学术指纹。

引用此