Skip to main navigation Skip to search Skip to main content

DRLFailureMonitor: A Dynamic Failure Monitoring Approach for Deep Reinforcement Learning System

  • Yi Cai*
  • , Xiaohui Wan
  • , Zhihao Liu
  • , Zheng Zheng
  • *Corresponding author for this work
  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Over the past decade, deep reinforcement learning (DRL) has seen increasing adoption in addressing various sequential decision-making tasks, such as autonomous driving and robotic control, demonstrating superior performance. However, as its application continues to broaden, the reliability of these DRL systems encounters significant challenges, particularly within safety-critical domains where any system failure could lead to catastrophic consequences. Currently, the assurance of reliability in DRL systems relies on testing techniques, which consist of offline solutions that uncover and address potential defects before deployment. In contrast to existing studies, this paper addresses the issue of online monitoring of DRL systems to detect and alert potential failures in advance, thereby facilitating the transition from automated to manual decision-making when necessary. Specifically, we model online monitoring of DRL systems as a multivariate time series classification problem and propose a novel failure monitoring approach, which is named DRLFailureMonitor. This method employs a temporal dynamic graph neural network to capture hidden spatiotemporal dependencies in the sequences of trajectories planned by the DRL systems. Extensive experiments across five benchmark DRL environments demonstrate that DRLFailureMonitor achieves an average failure detection accuracy of 98.2% and a recall of 100%. In addition, the proposed method offers a failure detection lead time ranging from 11 to 79 steps, indicating that it can detect failures in DRL systems well before the system actually fails and causes any losses. Consequently, this method holds significant importance for enhancing human-machine collaboration and improving the reliability of DRL systems in safety-critical areas.

Original languageEnglish
Title of host publicationProceedings - 2024 IEEE 35th International Symposium on Software Reliability Engineering, ISSRE 2024
PublisherIEEE Computer Society
Pages487-498
Number of pages12
ISBN (Electronic)9798350353884
DOIs
StatePublished - 2024
Event35th IEEE International Symposium on Software Reliability Engineering, ISSRE 2024 - Tsukuba, Japan
Duration: 28 Oct 202431 Oct 2024

Publication series

NameProceedings - International Symposium on Software Reliability Engineering, ISSRE
ISSN (Electronic)2332-6549

Conference

Conference35th IEEE International Symposium on Software Reliability Engineering, ISSRE 2024
Country/TerritoryJapan
CityTsukuba
Period28/10/2431/10/24

Keywords

  • Deep Reinforcement Learning
  • Failure Monitoring
  • Multivariate Time Series Classification

Fingerprint

Dive into the research topics of 'DRLFailureMonitor: A Dynamic Failure Monitoring Approach for Deep Reinforcement Learning System'. Together they form a unique fingerprint.

Cite this