首页|期刊导航|自动化学报|基于积分强化学习的学习感知型动态事件触发最优控制

基于积分强化学习的学习感知型动态事件触发最优控制OA

Scalable Dynamic Event-triggered Optimal Control for Nonlinear Systems via Integral Reinforcement Learning

中文摘要英文摘要

事件触发机制,尤其是动态事件触发机制,近年来在控制领域引起广泛关注,其核心挑战在于平衡控制性能与通信资源利用率.当该机制与学习系统结合时,这种平衡变得尤为关键,因为还需兼顾学习效率.针对具有未知动态的非线性连续时间系统,提出一种集成积分强化学习、最优控制与学习感知设计的新型动态事件触发最优控制方法,该方法采用仅含评价网络的自适应结构在线学习最优控制策略,并通过灵活配置的动态触发规则调控数据传输.其核心创新在于设计了一种学习感知型动态事件触发机制,该机制通过分析评价网络权值的历史变化趋势,构建学习感知参数,进而自适应地调整事件触发规则中的动态阈值参数.这使得系统能适宜地在学习关键期采用"繁忙采样"以保障控制与学习精度,在学习平稳期切换至"空闲采样"以节约通信与计算资源,从而实现控制性能、学习效率与资源消耗的有效平衡.理论分析严格证明了闭环系统的渐近稳定性和权值误差的一致最终有界性.最后,在一个基准非线性系统和一个单连杆机械臂系统进行了仿真验证与对比实验,结果表明与传统静态及动态事件触发方法相比,提出方法能以更少的通信代价获得相当甚至更优的学习与控制效果.

Event-triggered mechanisms,particularly dynamic ones,have garnered significant interest in the control community,with a key challenge being the balance between control performance and resource utilization.This bal-ance becomes even more crucial when integrating these mechanisms into learning systems,where learning efficiency plays a vital role.This paper presents a learning-based dynamic event-triggered framework that combines optimal control formulation,learning-aware design,and integral reinforcement learning,allowing the system to adapt the triggering process based on learning status and state changes.Using only partial knowledge of the dynamics,an op-timal control policy can be learned online via a critic neural network,with data transmission flexibly regulated by dynamic triggering rules.This enables the system to intelligently adopt"busy sampling"when the weight changes dramatically,and switch to"idle sampling"during smooth learning periods to save communication/computational resources,thereby achieving an effective balance between control performance,learning efficiency,and resource con-sumption.Theoretical analysis rigorously proves the asymptotic stability of closed-loop systems and the uniform ul-timate boundedness of weight errors.Finally,the proposed method is comparatively verified on a benchmark nonlin-ear system and a single-link robotic arm system,indicating that it can achieve comparable or even better learning and control effects with less communication cost.

王珂;许振钰;张俊楠;穆朝絮

天津大学电气自动化与信息工程学院 天津 300072天津大学电气自动化与信息工程学院 天津 300072天津大学电气自动化与信息工程学院 天津 300072天津大学电气自动化与信息工程学院 天津 300072

动态事件触发机制自适应动态规划积分强化学习最优控制学习感知设计

dynamic event-triggered mechanismsadaptive dynamic programmingintegral reinforcement learningoptimal controllearning-aware design

《自动化学报》 2026 (6)

1221-1233,13

国家自然科学基金(62503356,62333016),中国高校产学研创新基金(2024ZY009)资助 Supported by National Natural Science Foundation of China(62503356,62333016)and China Higher Education Institution In-dustry-University-Research Innovation Fund(2024ZY009)

10.16383/j.aas.c250620

评论