首页|期刊导航|无线电工程|基于Off-policy Q-学习的时延系统线性二次型跟踪控制算法

基于Off-policy Q-学习的时延系统线性二次型跟踪控制算法OA

Linear Quadratic Tracking Control Algorithm for Time-delay Systems Based on Off-policy Q-learning

中文摘要英文摘要

对被控系统数学模型参数未知的线性离散时间系统,同时考虑工业过程中数据存在控制输入时间延时的问题,提出一种数据驱动算法,解决时延系统线性二次型跟踪(Linear Quadratic Tracking,LQT)控制问题.通过对时延系统控制问题的描述,构建了基于模型驱动的强化学习算法框架,在此基础上为了避免使用数学模型参数和状态数据信息,引入Smith预估器,提出了基于On-policy Q-学习的时延系统LQT控制算法.考虑On-policy Q-学习算法中探测噪声对学习结果的影响,进一步采用Off-policy算法解决时延系统LQT控制问题.在此基础上改进Q-学习算法中使用的Bellman方程,提出数据驱动的Off-policy Q-学习算法,该算法不受探测噪声的影响,求解得到的解是无偏差的.理论分析和仿真实验表明,在避免依赖系统数学模型参数和状态数据的前提下,有效实现了时延系统的跟踪控制.

A data-driven algorithm is proposed to solve the Linear Quadratic Tracking(LQT)control problem for linear discrete-time systems with unknown model parameters,which also addresses the issue of control input time delays commonly encountered in industrial processes.Through the characterization of control problems in time-delay systems,a model-driven reinforcement learning framework is constructed,based on which a Smith predictor is introduced to avoid using the mathematical model parameters and state data,and a linear quadratic tracking control algorithm for time-delay systems is proposed based on On-policy Q-learning.Considering the impact of exploration noise on the learning results in the On-policy Q-learning algorithm,an Off-policy algorithm is further adopted to solve the linear quadratic tracking control problem for time-delay systems.On this basis,the Bellman equation used in the Q-learning algorithm is improved,and a data-driven Off-policy Q-learning algorithm is presented,which remains unaffected by exploration noise and provides unbiased solutions.Theoretical analysis and simulation experiments demonstrate that tracking control for time-delay systems is effectively achieved without reliance on system mathematical model parameters or state data.

刘文;蔚保国;郝菁;王卿

卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081

信息技术与安全科学

时延系统强化学习Off-policy数据驱动输出反馈

time-delay systemsreinforcement learningOff-policydata-drivenoutput feedback

《无线电工程》 2026 (1)

166-176,11

河北省创新能力提升计划(24460801D)Innovation Capability Promotion Plan of Hebei Province(24460801D)

10.3969/j.issn.1003-3106.2026.01.018

评论