Skip to main navigation Skip to search Skip to main content

Self-Imitation Learning for Robot Tasks with Sparse and Delayed Rewards

  • Beihang University
  • Nanjing Research Institute of Electronics Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The application of reinforcement learning (RL) in robotic control is still limited in the environments with sparse and delayed rewards. In this paper, we propose a practical self-imitation learning method named Self-Imitation Learning with Constant Reward (SILCR). Instead of requiring hand-defined immediate rewards from environments, our method assigns the immediate rewards at each timestep with constant values according to their final episodic rewards. In this way, even if the dense rewards from environments are unavailable, every action taken by the agents would be guided properly. We demonstrate the effectiveness of our method in some challenging continuous robotics control tasks in MuJoCo simulation and the results show that our method significantly outperforms the alternative methods in tasks with sparse and delayed rewards. Even compared with alternatives with dense rewards available, our method achieves competitive performance. The ablation experiments also show the stability and reproducibility of our method.

Original languageEnglish
Title of host publication2021 IEEE International Conference on Mechatronics and Automation, ICMA 2021
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages477-482
Number of pages6
ISBN (Electronic)9781665441001
DOIs
StatePublished - 8 Aug 2021
Event18th IEEE International Conference on Mechatronics and Automation, ICMA 2021 - Takamatsu, Japan
Duration: 8 Aug 202111 Aug 2021

Publication series

Name2021 IEEE International Conference on Mechatronics and Automation, ICMA 2021

Conference

Conference18th IEEE International Conference on Mechatronics and Automation, ICMA 2021
Country/TerritoryJapan
CityTakamatsu
Period8/08/2111/08/21

Keywords

  • Rewards delay
  • Robot control
  • SIL

Fingerprint

Dive into the research topics of 'Self-Imitation Learning for Robot Tasks with Sparse and Delayed Rewards'. Together they form a unique fingerprint.

Cite this