TY - JOUR
T1 - Role-Policy Enhanced Collaborative Task Learning in Multiagent Systems
AU - Fu, Hang
AU - Wang, Jingjing
AU - Wang, Ye
AU - Ren, Pengfei
AU - Chen, Jianrui
AU - Chen, Philip
N1 - Publisher Copyright:
© 2013 IEEE.
PY - 2026
Y1 - 2026
N2 - Current mainstream multiagent reinforcement learning (MARL) algorithms primarily focus on acquiring the global maximum reward throughout the entire training process, from the initial to the final stage. Whereas directly pursuing the global maximum return tends to be inefficient, particularly in environments with sparse rewards or the large-scale multiagent system. To address these challenges, previous algorithms have been developed to maintain individual policies to guide global training. Nevertheless, these approaches generally neglect either efficiency or the potential for local collaboration at the early stage of training. In this article, we propose the role-policy enhanced global policy (RPEGP) algorithm, which integrates the concept of distinct roles within the actor–critic-based MARL framework. RPEGP simultaneously considers both collaborative behaviors among agents and efficient global policy training. Specifically, RPEGP exploits the similarities among agents to assign distinct roles, training role-policies and the global policy concurrently. Through the initialization and enhancement of the role-policies, the global policy is trained more efficiently and effectively. Empirical experiments are conducted in well-known cooperative multiagent environments, including StarCraft II micromanagement (SMAC) and multiagent particle environment (MPE). Experimental results demonstrate that RPEGP outperforms baseline algorithms across various evaluation metrics and training efficiency, confirming its ability to address complex cooperative tasks generically and efficiently.
AB - Current mainstream multiagent reinforcement learning (MARL) algorithms primarily focus on acquiring the global maximum reward throughout the entire training process, from the initial to the final stage. Whereas directly pursuing the global maximum return tends to be inefficient, particularly in environments with sparse rewards or the large-scale multiagent system. To address these challenges, previous algorithms have been developed to maintain individual policies to guide global training. Nevertheless, these approaches generally neglect either efficiency or the potential for local collaboration at the early stage of training. In this article, we propose the role-policy enhanced global policy (RPEGP) algorithm, which integrates the concept of distinct roles within the actor–critic-based MARL framework. RPEGP simultaneously considers both collaborative behaviors among agents and efficient global policy training. Specifically, RPEGP exploits the similarities among agents to assign distinct roles, training role-policies and the global policy concurrently. Through the initialization and enhancement of the role-policies, the global policy is trained more efficiently and effectively. Empirical experiments are conducted in well-known cooperative multiagent environments, including StarCraft II micromanagement (SMAC) and multiagent particle environment (MPE). Experimental results demonstrate that RPEGP outperforms baseline algorithms across various evaluation metrics and training efficiency, confirming its ability to address complex cooperative tasks generically and efficiently.
KW - Actor–critic
KW - StarCraft II
KW - cooperative multiagent systems
KW - multiagent reinforcement learning (MARL)
KW - role-oriented learning
UR - https://www.scopus.com/pages/publications/105020758680
U2 - 10.1109/TSMC.2025.3623649
DO - 10.1109/TSMC.2025.3623649
M3 - 文章
AN - SCOPUS:105020758680
SN - 2168-2216
VL - 56
SP - 267
EP - 278
JO - IEEE Transactions on Systems, Man, and Cybernetics: Systems
JF - IEEE Transactions on Systems, Man, and Cybernetics: Systems
IS - 1
ER -