首页|期刊导航|山东电力技术|一种基于改进TD3算法的超临界机组协调控制方法

一种基于改进TD3算法的超临界机组协调控制方法OA

A Coordinated Control Strategy for Supercritical Units Based on the Improved TD3 Algorithm

中文摘要英文摘要

超临界机组机炉耦合严重、动态响应差异大,传统比例积分微分(proportional integral differential,PID)协调控制难以兼顾主蒸汽压力稳定与快速负荷响应.为此,提出一种基于双延迟深度确定性策略梯度算法(twin delayed deep deterministic policy gradient,TD3)的控制策略,通过引入双Critic评估网络与延迟策略更新,有效改善深度确定性策略梯度(deep deterministic policy gradient,DDPG)算法存在的训练不稳定问题,进而提升协调控制系统控制品质.首先,运用改进海鸥优化算法对机组协调系统进行模型辨识,建立被控对象的数学模型.然后,基于动作增强策略(action-enhancing strategies,AES)改进TD3算法的初期探索机制,加快算法收敛速度,并设计多维状态信息和多目标奖励函数,构建以TD3智能体为核心的控制框架,通过TD3算法生成最优的燃料控制指令,提升炉侧主蒸汽压力响应速度与控制精度.最后,基于600 MW超临界机组的Simulink仿真验证表明,该策略有效提升了控制品质.

The coupling between the steam turbine and boiler in supercritical units is severe,and the dynamic response varies greatly.The traditional proportional integral differential(PID)coordinated control is difficult to balance the stability of the main steam pressure and rapid load response.To this end,a control strategy based on the twin delayed deep deterministic policy gradient(TD3)algorithm is proposed.Through the introduction of the double Critic evaluation network and the delayed policy update,The training instability issue of the deep deterministic policy gradient(DDPG)algorithm has been effectively addressed,which enhances the control quality of the coordinated control system.Firstly,the improved seagull optimization algorithm is applied to conduct model identification of the unit coordination system,and establishes mathematical models of the controlled object;Secondly,the initial exploration mechanism of TD3 algorithm is improved based on action-enhancing strategies(AES)to speed up the convergence speed of the algorithm.The multi-dimensional state information and multi-objective reward function are designed,and a control framework centered on the TD3 agent is constructed.The optimal fuel control command is generated through TD3 algorithm to improve the response speed and control accuracy of the main steam pressure on the boiler side.Finally,the simulink simulation based on 600 MW supercritical unit shows that this strategy can effectively improve the control quality.

胥晓晶;何同祥;郑晓刚

华北电力大学控制与计算机工程学院,河北 保定 071000华北电力大学控制与计算机工程学院,河北 保定 071000华北电力大学控制与计算机工程学院,河北 保定 071000

信息技术与安全科学

TD3海鸥优化算法协调控制系统多维状态信息多目标奖励函数

TD3seagull optimization algorithmcoordinated control systemsmulti-dimensional state informationmulti-objective reward function

《山东电力技术》 2026 (5)

112-120,9

10.20097/j.cnki.issn1007-9904.250697

评论