TY - JOUR
T1 - Two-stage Distributed Generators Optimization Based on Deep Reinforcement Learning With Parameter Sharing
AU - Gao, Fang
AU - Yao, Haotian
AU - Gao, Qing
AU - Yin, Linfei
AU - Cai, Yunxiang
AU - Jin, Yan
AU - Pan, Yu
N1 - Publisher Copyright:
©2025 Chin.Soc.for Elec.Eng.
PY - 2025/10/5
Y1 - 2025/10/5
N2 - As renewable energy sources such as solar and wind are increasingly integrated into the grid at high proportions, the optimization and scheduling of distributed power generation face challenges due to frequent changes in system topology, affecting the stability and economic operation of the distribution network. Existing methods, designed for systems with fixed topology, rely on precise models and are time-consuming, making real-time control difficult. Current deep reinforcement learning approaches struggle to balance distributed training and mixed discrete-continuous action spaces. This study introduces a distributed power optimization strategy based on multi-agent deep reinforcement learning with two stages and parameter sharing. Initially, the problem is vertically decoupled, constructing a dynamic distribution network reconfiguration model with distributed generation using mixed-integer second-order cone programming to determine the topology. Subsequently, the distribution network environment is horizontally decoupled into several regions. In the second stage, a centralized training with decentralized execution framework that incorporates parameter sharing is proposed. This framework incorporates a multi-agent prioritized double-delay deep deterministic policy gradient algorithm with a priority experience replay mechanism. Topology information is embedded into the distribution network environment, mapped to agents through power flow calculations to minimize network active power loss in the optimization scheduling model. Case studies demonstrate that the proposed algorithm, by considering changes in the distribution network topology and enhancing learning efficiency through strategy and experience sharing among agents, as well as priority experience replay, meets the efficiency requirements of real-time online decision-making and shows superior voltage stability and loss reduction performance compared to other strategies.
AB - As renewable energy sources such as solar and wind are increasingly integrated into the grid at high proportions, the optimization and scheduling of distributed power generation face challenges due to frequent changes in system topology, affecting the stability and economic operation of the distribution network. Existing methods, designed for systems with fixed topology, rely on precise models and are time-consuming, making real-time control difficult. Current deep reinforcement learning approaches struggle to balance distributed training and mixed discrete-continuous action spaces. This study introduces a distributed power optimization strategy based on multi-agent deep reinforcement learning with two stages and parameter sharing. Initially, the problem is vertically decoupled, constructing a dynamic distribution network reconfiguration model with distributed generation using mixed-integer second-order cone programming to determine the topology. Subsequently, the distribution network environment is horizontally decoupled into several regions. In the second stage, a centralized training with decentralized execution framework that incorporates parameter sharing is proposed. This framework incorporates a multi-agent prioritized double-delay deep deterministic policy gradient algorithm with a priority experience replay mechanism. Topology information is embedded into the distribution network environment, mapped to agents through power flow calculations to minimize network active power loss in the optimization scheduling model. Case studies demonstrate that the proposed algorithm, by considering changes in the distribution network topology and enhancing learning efficiency through strategy and experience sharing among agents, as well as priority experience replay, meets the efficiency requirements of real-time online decision-making and shows superior voltage stability and loss reduction performance compared to other strategies.
KW - deep reinforcement learning
KW - distributed generation optimization scheduling
KW - dynamic reconfiguration
KW - parameter sharing
KW - topological changes
UR - https://www.scopus.com/pages/publications/105027206436
U2 - 10.13334/j.0258-8013.pcsee.240635
DO - 10.13334/j.0258-8013.pcsee.240635
M3 - 文章
AN - SCOPUS:105027206436
SN - 0258-8013
VL - 45
SP - 7493
EP - 7509
JO - Zhongguo Dianji Gongcheng Xuebao/Proceedings of the Chinese Society of Electrical Engineering
JF - Zhongguo Dianji Gongcheng Xuebao/Proceedings of the Chinese Society of Electrical Engineering
IS - 19
ER -