Skip to main navigation Skip to search Skip to main content

Human Interaction Reinforcement Learning With Disturbance for Muti-Agent Systems Consensus Control

  • Beihang University

Research output: Contribution to journalConference articlepeer-review

Abstract

To address the critical challenges of multi-agent reinforcement learning under human interactive inputs, this paper proposes a novel robust and interpretable reinforcement learning framework. First, a robustness optimization module based on an enhanced actor-critic architecture is designed to effectively mitigate disturbances arising from operational errors and subjective biases in human inputs, thereby improving system fault tolerance while ensuring policy convergence. Second, Lyapunov stability theory is innovatively incorporated into the policy optimization process, providing rigorous mathematical proofs of stability and endowing agent behaviors with white-box interpretability. Finally, by leveraging reinforcement learning design, the approach overcomes the reliance of traditional methods on continuous incentive signals, significantly enhancing algorithmic applicability in open and dynamic environments. Extensive comparative experiments validate the superior performance of the proposed method in training efficiency, policy robustness, and interpretability, offering a reliable solution for the deployment of human-machine collaborative systems in complex scenarios.

Original languageEnglish
Pages (from-to)140-145
Number of pages6
JournalInternational Conference on Robotics and Automation Sciences, ICRAS
Issue number2025
DOIs
StatePublished - 2025
Event9th International Conference on Robotics and Automation Sciences, ICRAS 2025 - Osaka, Japan
Duration: 27 Jun 202529 Jun 2025

Keywords

  • Consensus tracking
  • Disturbance input
  • Human interacted
  • Multi-agent
  • Reinforcement learning

Fingerprint

Dive into the research topics of 'Human Interaction Reinforcement Learning With Disturbance for Muti-Agent Systems Consensus Control'. Together they form a unique fingerprint.

Cite this