弥合配电系统恢复调度仿真-现实间隙的两阶段数据-机理融合优化架构OA
Two-stage Data-mechanism Integrated Optimization Framework to Bridge the Simulation-to-reality Gap of Distribution System Restoration Dispatch
配电系统的灾后恢复调度是提升配电系统韧性、保障经济发展与社会民生的重要环节.现有研究已初步探索以深度强化学习为代表的数据驱动方法在配电系统恢复调度中的应用模式,但当训练调度智能体所用的仿真场景与实际灾后场景差异较大时,智能体的实际决策性能将产生严重下降,该现象被称为仿真-现实间隙.针对该问题,提出一种具有动态泛化性的配电系统灾后恢复调度方法,一方面构建灾前-灾中两阶段混合约束马尔可夫决策过程模型,灾中基于已知的故障信息对调度策略进行深度调优,规避故障位置偏差引起的强仿真-现实间隙;另一方面设计机理引导的恢复调度智能体训练方法提高训练效率,进而允许灾中2阶段训练通过适当扩大场景集以尽可能覆盖各类修复用时与用户侧源荷功率,进一步缓解仿真-现实间隙.基于双变电站66 节点测试系统的算例测试,验证所提方法较传统深度强化学习方法有更优的训练效率、收敛性与泛化性.
The post-disaster restoration dispatch of distribution systems is a crucial element in enhancing its resilience and ensuring economic development and societal welfare.Existing studies have preliminarily explored the application of data-driven methods,represented by deep reinforcement learning,in distribution system restoration dispatch.However,when the simulation scenario used to train the dispatch agent differs significantly from the real scenario,the decision-making performance of the agent will produce a serious degradation,which is known as the simulation-to-reality gap.To address this issue,this paper proposes a distribution system post-disaster restoration dispatch method with dynamic generalization capabilities.On the one hand,a two-stage(before and during the disaster)hybrid-constrained Markov decision process formulation is developed,where the dispatch policy is further fine-tuned based on known fault information during the disaster to bridge the simulation-to-reality gap caused by uncertain fault locations.On the other hand,this paper designs a mechanism-guided restoration dispatch agent training approach to improve training efficiency,thereby enabling the second-stage fine-tuning to cover an expanded scenario set and mitigate the simulation-to-reality gap induced by uncertain repair durations and prosumer-side power.Test results in case studies based on a dual-substation 66-bus test system validate that the proposed method outperforms traditional deep learning methods in terms of training efficiency,convergence,and generalization capability.
吴奕之;刘浏;康重庆;叶宇剑;胡健雄;周东华;朱瑨;张新松;杨春祥;马彦宏;赵沛霖
东南大学电气工程学院,江苏省 南京市 210096深圳市腾讯计算机系统有限公司,广东省 深圳市 518063清华大学电机工程与应用电子技术系,北京市 海淀区 100084东南大学电气工程学院,江苏省 南京市 210096温州大学电气与电子工程学院,浙江省 温州市 325035东南大学自动化学院,江苏省 南京市 210096东南大学土木工程学院,江苏省 南京市 211189南通大学电气与自动化学院,江苏省 南通市 226019国网甘肃省电力公司电力调度中心,甘肃省 兰州市 730030国网甘肃省电力公司电力调度中心,甘肃省 兰州市 730030上海交通大学人工智能学院,上海市 闵行区 200000
信息技术与安全科学
配电系统恢复调度分布式资源多阶段随机规划仿真-现实间隙数据机理融合
distribution system restoration dispatchdistributed energy resourcesmulti-stage stochastic programmingsimulation-to-reality gapdata-mechanism integration
《中国电机工程学报》 2026 (15)
6214-6226,中插6,14
国家自然科学基金项目(52477083,52207082,52507139)腾讯高校合作项目(Tencent JR2025TEG001).Project Supported by National Natural Science Foundation of China(52477083,52207082,52507139)Tencent-University Collaboration Program(Tencent JR2025TEG001).
评论