基于Off-policy Q-学习的时延系统线性二次型跟踪控制算法OA
Linear Quadratic Tracking Control Algorithm for Time-delay Systems Based on Off-policy Q-learning
对被控系统数学模型参数未知的线性离散时间系统,同时考虑工业过程中数据存在控制输入时间延时的问题,提出一种数据驱动算法,解决时延系统线性二次型跟踪(Linear Quadratic Tracking,LQT)控制问题.通过对时延系统控制问题的描述,构建了基于模型驱动的强化学习算法框架,在此基础上为了避免使用数学模型参数和状态数据信息,引入Smith预估器,提出了基于On-policy Q-学习的时延系统LQT控制算法.考虑On-policy Q-学习算法中探测噪声对学习结果的影响,进一步采用Off-policy算法解决时延系统LQT控制问题.在此基础上改进Q-学习算法中使用的Bellman方程,提出数据驱动的Off-policy Q-学习算法,该算法不受探测噪声的影响,求解得到的解是无偏差的.理论分析和仿真实验表明,在避免依赖系统数学模型参数和状态数据的前提下,有效实现了时延系统的跟踪控制.
A data-driven algorithm is proposed to solve the Linear Quadratic Tracking(LQT)control problem for linear discrete-time systems with unknown model parameters,which also addresses the issue of control input time delays commonly encountered in industrial processes.Through the characterization of control problems in time-delay systems,a model-driven reinforcement learning framework is constructed,based on which a Smith predictor is introduced to avoid using the mathematical model parameters and state data,and a linear quadratic tracking control algorithm for time-delay systems is proposed based on On-policy Q-learning.Considering the impact of exploration noise on the learning results in the On-policy Q-learning algorithm,an Off-policy algorithm is further adopted to solve the linear quadratic tracking control problem for time-delay systems.On this basis,the Bellman equation used in the Q-learning algorithm is improved,and a data-driven Off-policy Q-learning algorithm is presented,which remains unaffected by exploration noise and provides unbiased solutions.Theoretical analysis and simulation experiments demonstrate that tracking control for time-delay systems is effectively achieved without reliance on system mathematical model parameters or state data.
刘文;蔚保国;郝菁;王卿
卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081卫星导航系统与装备技术国家重点实验室,河北 石家庄 050081||中国电子科技集团公司第五十四研究所,河北 石家庄 050081
信息技术与安全科学
时延系统强化学习Off-policy数据驱动输出反馈
time-delay systemsreinforcement learningOff-policydata-drivenoutput feedback
《无线电工程》 2026 (1)
166-176,11
河北省创新能力提升计划(24460801D)Innovation Capability Promotion Plan of Hebei Province(24460801D)
评论