Skip to main navigation Skip to search Skip to main content

A Multi-sensing Input and Multi-constraint Reward Mechanism Based Deep Reinforcement Learning Method for Self-driving Policy Learning

  • Zhongli Wang*
  • , Hao Wang
  • , Xin Cui
  • , Chaochao Zheng
  • *Corresponding author for this work
  • Beijing Jiaotong University
  • Beijing Engineering Research Center of EMC and GNSS Technology for Rail Transportation
  • Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Planning and decision-making of autonomous driving is an active and challenging topic currently. Deep reinforcement learning-based approaches seek to solve the problem in an end-to-end manner, but generally require a large amount of sample data and confronted with high dimensionality of input data and complex models, which lead to slow convergence and cannot learn effectively with noisy data. Most of deep reinforcement learning-based approaches use a sample reward function. Due to the complicated and volatile traffic scenarios, these approaches cannot satisfy the driving policy requirement. To address the issues, a multi-sensing and multi-constraint reward function (MSMC-SAC) based deep reinforcement learning method is proposed. The inputs of the proposed method include front-view image, point cloud from LiDAR, as well as the bird's-eye view generated from the perception results. The multi-sensing input is first passed to an encoding network to obtain the representation in latent space and then forward to a SAC-based learning module. A multiple rewards function considering various constraints, such as the error of transverse-longitudinal distance and heading angle, smoothness, velocity, and the possibility of collision, is designed. The performance of the proposed method in different typical traffic scenarios is validated with CARLA [1]. The effects of multiple reward functions are compared. The simulation results show that the presented approach can learn the driving policies in many complex scenarios, such as straight ahead, passing the intersections, and making turning, and outperforms against the existing typical deep reinforcement learning methods.

Original languageEnglish
Title of host publicationIntelligent Robotics and Applications - 14th International Conference, ICIRA 2021, Proceedings
EditorsXin-Jun Liu, Zhenguo Nie, Jingjun Yu, Fugui Xie, Rui Song
PublisherSpringer Science and Business Media Deutschland GmbH
Pages691-701
Number of pages11
ISBN (Print)9783030890919
DOIs
StatePublished - 2021
Externally publishedYes
Event14th International Conference on Intelligent Robotics and Applications, ICIRA 2021 - Yantai, China
Duration: 22 Oct 202125 Oct 2021

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume13016 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference14th International Conference on Intelligent Robotics and Applications, ICIRA 2021
Country/TerritoryChina
CityYantai
Period22/10/2125/10/21

Keywords

  • CARLA
  • Deep reinforcement learning
  • Driving policy
  • Multi-reward functions

Fingerprint

Dive into the research topics of 'A Multi-sensing Input and Multi-constraint Reward Mechanism Based Deep Reinforcement Learning Method for Self-driving Policy Learning'. Together they form a unique fingerprint.

Cite this