摘要
Multi-agent reinforcement learning (MARL) has attracted more and more attention in recent years. It is now widely applied in various fields, including cyber physical systems, smart grid, finance, social network, and among others. The current researches on MARL mainly focus single-time scale, in which the agents have the same decision epoch. While in real applications, it is common that the agents make decisions by different frequencies. In addition, different agents may have separate roles in the system. In this paper, we propose a more general MARL framework by introducing multi-time scale of decision epochs. We assume that agents share information with their neighbors, including state, action, and reward. The global observability of state and action, which is a common assumption, is not required. We propose a decentralized Q-learning algorithm and a modified MADDPG algorithm to solve the problem. The main contributions of this paper are as follows. First, we formulate the multi-time scale multi-agent reinforcement learning (MTMARL) problem. This provides a general framework for the related systems and problems. Second, we provide a networked decentralized multi-time scale multi-agent Q-learning algorithm to solve the problem and prove its convergence. Third, we test the algorithm numerically. The results show that the proposed algorithm performs better than the previous QD-learning and is only slightly worse than the centralized algorithm.
| 源语言 | 英语 |
|---|---|
| 主期刊名 | 2020 59th IEEE Conference on Decision and Control, CDC 2020 |
| 出版商 | Institute of Electrical and Electronics Engineers Inc. |
| 页 | 578-584 |
| 页数 | 7 |
| ISBN(电子版) | 9781728174471 |
| DOI | |
| 出版状态 | 已出版 - 14 12月 2020 |
| 已对外发布 | 是 |
| 活动 | 59th IEEE Conference on Decision and Control, CDC 2020 - Virtual, Online, 韩国 期限: 14 12月 2020 → 18 12月 2020 |
丛书
| 姓名 | Proceedings of the IEEE Conference on Decision and Control |
|---|---|
| 卷 | 2020-December |
| ISSN(印刷版) | 0743-1546 |
| ISSN(电子版) | 2576-2370 |
会议
| 会议 | 59th IEEE Conference on Decision and Control, CDC 2020 |
|---|---|
| 国家/地区 | 韩国 |
| 市 | Virtual, Online |
| 时期 | 14/12/20 → 18/12/20 |
联合国可持续发展目标
此成果有助于实现下列可持续发展目标:
-
可持续发展目标 7 经济适用的清洁能源
学术指纹
探究 'Decentralized Multi-agent Reinforcement Learning with Multi-time Scale of Decision Epochs' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver