水库群发电调度多智能体深度强化学习方法OA
Multi-agent deep reinforcement learning method for cascade reservoir operation for power generation
水库群发电调度问题中,传统动态规划等方法存在计算效率低等瓶颈,制约其实际工程应用;单智能体深度强化学习虽能以端到端方式实现连续控制,却难以有效刻画上下游水库间的动态关联,易引发策略梯度失稳与训练振荡,且将梯级系统视为单一整体,可扩展性受限.为此,本文提出一种基于"集中式评价-分布式执行"框架的水库群发电调度多智能体深度强化学习方法:首先,将各水库建模为独立智能体,构建局部决策与全局信息共享的协同网络以提升训练稳定性;其次,设计融合发电效益目标与约束惩罚项的全局奖励函数,并引入Copula-Gibbs联合分布生成径流情景,增强对来水不确定性的适应能力;最后,通过敏感性分析与网格搜索优选神经网络结构、折扣因子及学习率等超参数组合.工程应用结果表明:在相同硬件与数据条件下,所提方法具有更快的收敛速度;在严格满足调度期末水位与运行安全约束的前提下,在4库和6库系统中实现6.5~8.5 ms的在线推理响应速度,较离散微分动态规划提升约2个数量级;在枯、平、丰三种典型来水情景下,年发电量偏差控制在0.65%以内,且期末水位控制精度良好.综上,多智能体协同机制可有效应对高维决策与系统耦合带来的挑战,提升调度策略的环境适应性与工程实用性,为大规模水库群高效调度提供了可靠方法支撑.
In the optimization of cascade reservoir operation for power generation,traditional methods such as dynamic programming suffer from bottlenecks like low computational efficiency,limiting their practical engineering applications.Although single-agent deep reinforcement learning(DRL)can achieve continuous control in an end-to-end manner,it struggles to effectively capture the dynamic couplings among cascade reservoirs,potentially leading to unstable policy gradients and training oscillations.Moreover,treating the cascade reservoir system as a single entity limits scalability.To address these issues,this paper proposes a multi-agent deep reinforcement learning(MADRL)method for cascade reservoir operation based on the"Centralized Training and Decentralized Execution"framework.Each reservoir is treated as an independent agent,and a collaborative network integrating local decision-making with global information sharing is constructed to enhance training stability.A global reward function is designed,combin-ing power generation benefit with constraint penalty terms.Furthermore,the Copula-Gibbs joint distribution is intro-duced to generate runoff scenarios,improving adaptability to inflow uncertainty.Finally,hyperparameter combina-tions,including network architecture,discount factor,and learning rate,are tuned through sensitivity analysis and grid search.Engineering application results show that,under identical hardware and data conditions,the proposed method can achieve faster convergence.While strictly adhering to end-of-period water level and operational safety constraints,it achieves online inference times of 6.5 to 8.5 ms in 4 and 6 reservoirs cascaded system,which is approximately two orders of magnitude faster than discrete differential dynamic programming(DDDP).Under dry,normal,and wet typical inflow scenarios,the annual power generation deviation is controlled within 0.65%,and the end-of-period water level control accuracy is satisfactory.In conclusion,the multi-agent collaborative mechanism effectively addresses challenges posed by high-dimensional decision-making and system coupling,enhancing the environmental adaptability and engineering practicality of operational strategies.This provides reliable methodologi-cal support for efficient operation of large-scale cascade reservoirs.
冯仲恺;罗涛;关铁生;牛文静
河海大学 水灾害防御全国重点实验室,江苏 南京 210098||河海大学 水文水资源学院,江苏 南京 210098河海大学 水灾害防御全国重点实验室,江苏 南京 210098||河海大学 水文水资源学院,江苏 南京 210098水利部南京水利水文自动化研究所,江苏 南京 210012长江水利委员会水文局,湖北 武汉 430010
建筑与水利
水库群调度多智能体强化学习深度神经网络实时决策推理
cascade reservoir operationmulti-agent reinforcement learningdeep neural networksreal-time decision-making
《水利学报》 2026 (7)
985-1003,19
国家自然科学基金项目(52394234,52441901,52379009)江苏省自然科学基金优秀青年基金项目(BK20240189)北京江河水利发展基金会—水利青年科技英才项目(JHYC202310)江苏省科技智库计划项目(JSKX0225047)
评论