Abstract
Multi-agent reinforcement learning (MARL) has attracted more and more attention in recent years. It is now widely applied in various fields, including cyber physical systems, smart grid, finance, social network, and among others. The current researches on MARL mainly focus single-time scale, in which the agents have the same decision epoch. While in real applications, it is common that the agents make decisions by different frequencies. In addition, different agents may have separate roles in the system. In this paper, we propose a more general MARL framework by introducing multi-time scale of decision epochs. We assume that agents share information with their neighbors, including state, action, and reward. The global observability of state and action, which is a common assumption, is not required. We propose a decentralized Q-learning algorithm and a modified MADDPG algorithm to solve the problem. The main contributions of this paper are as follows. First, we formulate the multi-time scale multi-agent reinforcement learning (MTMARL) problem. This provides a general framework for the related systems and problems. Second, we provide a networked decentralized multi-time scale multi-agent Q-learning algorithm to solve the problem and prove its convergence. Third, we test the algorithm numerically. The results show that the proposed algorithm performs better than the previous QD-learning and is only slightly worse than the centralized algorithm.
| Original language | English |
|---|---|
| Title of host publication | 2020 59th IEEE Conference on Decision and Control, CDC 2020 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 578-584 |
| Number of pages | 7 |
| ISBN (Electronic) | 9781728174471 |
| DOIs | |
| State | Published - 14 Dec 2020 |
| Externally published | Yes |
| Event | 59th IEEE Conference on Decision and Control, CDC 2020 - Virtual, Online, Korea, Republic of Duration: 14 Dec 2020 → 18 Dec 2020 |
Publication series
| Name | Proceedings of the IEEE Conference on Decision and Control |
|---|---|
| Volume | 2020-December |
| ISSN (Print) | 0743-1546 |
| ISSN (Electronic) | 2576-2370 |
Conference
| Conference | 59th IEEE Conference on Decision and Control, CDC 2020 |
|---|---|
| Country/Territory | Korea, Republic of |
| City | Virtual, Online |
| Period | 14/12/20 → 18/12/20 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 7 Affordable and Clean Energy
Keywords
- Multi-time scale
- multi-agent
- multiple decision epochs
- reinforcement learning
Fingerprint
Dive into the research topics of 'Decentralized Multi-agent Reinforcement Learning with Multi-time Scale of Decision Epochs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver