TY - JOUR
T1 - Accelerating wargaming reinforcement learning by dynamic multi-demonstrator ensemble
AU - Dong, Liwei
AU - Li, Ni
AU - Yuan, Haitao
AU - Gong, Guanghong
N1 - Publisher Copyright:
© 2023 Elsevier Inc.
PY - 2023/11
Y1 - 2023/11
N2 - Deep Reinforcement Learning (DRL) has become a promising technique to deal with tough wargaming decision-making problems. However, DRL suffers an inherent problem of low learning efficiency and it often requires massive cost of training steps, which may be alleviated with expert demonstrations in wargaming domains. Most learning methods with demonstrations generally treat the demonstration data from different expert demonstrators without distinction. Besides, a more appropriate and effective mechanism is highly needed to control sampling balance of expert-generated demonstration samples and agent-generated interaction ones. To tackle the two issues, this work proposes an improved approach to leverage expert demonstrations to further accelerate DRL. It innovatively extracts inherent diversity in multiple demonstrators by pre-training agents individually from multiple demonstration sources, thereby producing a strong and initial ensemble model. In addition, a novel technique to evaluate the learning importance of each demonstrator is designed to dynamically tune sampling ratios of learning data in a more adaptive and effective manner. Through the evaluation on several classic game tasks and a typical wargaming scenario, our method shows superior performance over several state-of-the-art methods and significantly raises DRL's efficiency for typical wargaming decision-making applications.
AB - Deep Reinforcement Learning (DRL) has become a promising technique to deal with tough wargaming decision-making problems. However, DRL suffers an inherent problem of low learning efficiency and it often requires massive cost of training steps, which may be alleviated with expert demonstrations in wargaming domains. Most learning methods with demonstrations generally treat the demonstration data from different expert demonstrators without distinction. Besides, a more appropriate and effective mechanism is highly needed to control sampling balance of expert-generated demonstration samples and agent-generated interaction ones. To tackle the two issues, this work proposes an improved approach to leverage expert demonstrations to further accelerate DRL. It innovatively extracts inherent diversity in multiple demonstrators by pre-training agents individually from multiple demonstration sources, thereby producing a strong and initial ensemble model. In addition, a novel technique to evaluate the learning importance of each demonstrator is designed to dynamically tune sampling ratios of learning data in a more adaptive and effective manner. Through the evaluation on several classic game tasks and a typical wargaming scenario, our method shows superior performance over several state-of-the-art methods and significantly raises DRL's efficiency for typical wargaming decision-making applications.
KW - Decision-making
KW - Expert demonstrations
KW - Reinforcement learning
KW - Wargaming
UR - https://www.scopus.com/pages/publications/85172466827
U2 - 10.1016/j.ins.2023.119534
DO - 10.1016/j.ins.2023.119534
M3 - 文章
AN - SCOPUS:85172466827
SN - 0020-0255
VL - 648
JO - Information Sciences
JF - Information Sciences
M1 - 119534
ER -