多智能体强化学习驱动的园区产消者P2P能源交易策略OA
Multi-Agent Reinforcement Learning-Driven P2P Energy Trading Strategy for Community Prosumers
[目的]针对点对点(peer-to-peer,P2P)园区能源交易场景中产消者在资源配置与用能行为的差异,以及传统基于模型的方法在不确定环境下适应性不足的问题,提出一种兼具可扩展性与隐私保护能力的多智能体强化学习方法.[方法]首先,构建3类代表性异质化产消者模型;其次,建立基于中端市场费率定价机制的园区能源交易模型,并引入灵活性激励机制;最后,将产消者的能源交易决策问题构建成部分可观测马尔可夫决策过程,并提出一种动态均值场软策略-评价(dynamic mean-field soft actor-critic,DMF-SAC)算法,进而求解产消者的能源管理策略.[结果]仿真表明,该方法在收敛性能、计算开销及运行成本等方面均优于对比方法,并能有效提升分布式能源的就地消纳能力与削峰填谷水平.[结论]所提方法在兼顾隐私保护与系统扩展性的同时,有效提升了异质化产消者协同优化的效率与效益,对园区型市场能源交易与能量管理具有重要意义.
[Objective]To address the variations in resource allocation and energy consumption behavior among prosumers in peer-to-peer(P2P)community energy trading scenarios,as well as the limited adaptability of traditional model-based methods in uncertain environments,this paper proposes a multi-agent reinforcement learning method that features both scalability and privacy protection capabilities.[Methods]First,three representative types of heterogeneous prosumer models are constructed.Second,a community energy trading model based on the mid-market rate pricing mechanism is established,and a flexibility incentive mechanism is introduced.Finally,the energy trading decision-making problem of prosumers is formulated as a partially observable Markov decision process,and a soft actor-critic algorithm based on dynamic mean-field(DMF-SAC)approximation is proposed to solve the energy management strategies of prosumers.[Results]Simulation results demonstrate that the proposed method outperforms baseline methods in terms of convergence performance,computational overhead,and operating costs.It also effectively improves the local consumption of distributed energy and enhances peak-shaving and valley-filling capabilities.[Conclusions]The proposed method effectively improves the efficiency and economic benefits of collaborative optimization for heterogeneous prosumers while balancing privacy protection and system scalability,which holds significant value for energy trading and management in community-based markets.
陈桂力;陈丹红;文宏武;郑勇伟;麦亮;冯霞山;钟麒深;胡一鸣;郑杰辉
广东电网有限责任公司湛江供电局,广东省 湛江市 524002广东电网有限责任公司湛江供电局,广东省 湛江市 524002广东电网有限责任公司湛江供电局,广东省 湛江市 524002广东电网有限责任公司湛江供电局,广东省 湛江市 524002广东电网有限责任公司湛江供电局,广东省 湛江市 524002广东电网有限责任公司湛江供电局,广东省 湛江市 524002广东电网有限责任公司湛江供电局,广东省 湛江市 524002华南理工大学电力学院,广州市 510641华南理工大学电力学院,广州市 510641
信息技术与安全科学
点对点(P2P)能源交易能源管理多智能体强化学习隐私保护动态均值场软策略-评价(DMF-SAC)算法
peer-to-peer(P2P)energy tradingenergy managementmulti-agent reinforcement learningprivacy preservationdynamic mean-field soft actor-critic(DMF-SAC)algorithm
《电力建设》 2026 (8)
14-25,12
This work is supported by National Natural Science Foundation of China(No.52477097). 国家自然科学基金项目(52477097)
评论