TY - GEN
T1 - Robotic grasping training using deep reinforcement learning with policy guidance mechanism
AU - Yao, Junying
AU - Liu, Yongkui
AU - Lin, Tingyu
AU - Ping, Xubin
AU - Xu, He
AU - Wang, Wenxiao
AU - Xiao, Yingying
AU - Zhang, Lin
AU - Wang, Lihui
N1 - Publisher Copyright:
Copyright © 2021 by ASME
PY - 2021
Y1 - 2021
N2 - For the past few years, training robots to enable them to learn various manipulative skills using deep reinforcement learning (DRL) has arisen wide attention. However, large search space, low sample quality, and difficulties in network convergence pose great challenges to robot training. This paper deals with assembly-oriented robot grasping training and proposes a DRL algorithm with a new mechanism, namely, policy guidance mechanism (PGM). PGM can effectively transform useless or low-quality samples to useful or high-quality ones. Based on the improved Deep Q Network algorithm, an end-to-end policy model that takes images as input and outputs actions is established. Through continuous interactions with the environment, robots are able to learn how to optimally grasp objects according to the location of maximum Q value. A number of experiments for different scenarios using simulations and physical robots are conducted. Results indicate that the proposed DRL algorithm with PGM is effective in increasing the success rate of robot grasping, and moreover, is robust to changes of environment and objects.
AB - For the past few years, training robots to enable them to learn various manipulative skills using deep reinforcement learning (DRL) has arisen wide attention. However, large search space, low sample quality, and difficulties in network convergence pose great challenges to robot training. This paper deals with assembly-oriented robot grasping training and proposes a DRL algorithm with a new mechanism, namely, policy guidance mechanism (PGM). PGM can effectively transform useless or low-quality samples to useful or high-quality ones. Based on the improved Deep Q Network algorithm, an end-to-end policy model that takes images as input and outputs actions is established. Through continuous interactions with the environment, robots are able to learn how to optimally grasp objects according to the location of maximum Q value. A number of experiments for different scenarios using simulations and physical robots are conducted. Results indicate that the proposed DRL algorithm with PGM is effective in increasing the success rate of robot grasping, and moreover, is robust to changes of environment and objects.
KW - DRL
KW - Industrial robot training
KW - PGM
UR - https://www.scopus.com/pages/publications/85112507406
U2 - 10.1115/MSEC2021-63974
DO - 10.1115/MSEC2021-63974
M3 - 会议稿件
AN - SCOPUS:85112507406
T3 - Proceedings of the ASME 2021 16th International Manufacturing Science and Engineering Conference, MSEC 2021
BT - Manufacturing Processes; Manufacturing Systems; Nano/Micro/Meso Manufacturing; Quality and Reliability
PB - American Society of Mechanical Engineers
T2 - ASME 2021 16th International Manufacturing Science and Engineering Conference, MSEC 2021
Y2 - 21 June 2021 through 25 June 2021
ER -