空中机器人悬停控制的强化学习算法OA
Reinforcement learning for aerial robot hovering control
空中机器人作为空中作业的飞行载体,对平飞性能有较高要求,其携带的装备以及作业时的反作用力影响也要求空中机器人具有较强的稳定位姿能力.本文以提高强化学习效率为核心,研究空中机器人的悬停控制问题.首先,基于actor-critic算法提出一种增加监督器的策略梯度算法;其次,提出动态监督策略并基于PID控制器设计了一个动态监督器;然后,进行模拟仿真,分别利用基于行为价值Q的actor-critic算法(QAC)和深度确定性策略梯度算法(DDPG)作为对比算法进行机器人的悬停控制学习;最后,将算法部署到实际环境.仿真和实验结果显示,所设计的强化学习算法能够实现空中机器人的悬停控制自学习,并且与QAC和DDPG相比具有更高的学习效率、更快的收敛速度和更平稳的悬停效果.
As flying carriers for aerial operations,aerial robots have high requirements for level flight performance.The equipment they carry and the impact of reaction forces during operation also require aerial robots to have strong stable pose capabilities.This article focuses on hovering control for aerial robots,with an emphasis on improving reinforcement learning efficiency for this task.Firstly,the actor-critic algorithm based on behavioural value Q(QAC)and the deep deter-ministic policy gradient algorithm(DDPG)are used for robot hovering control learning.Subsequently,a policy gradient algorithm with added supervisors was proposed,using a PID controller as the dynamic monitor.Finally,simulation was conducted and the algorithm was deployed to the actual environment.Simulation and experimental results show that the designed reinforcement learning algorithm can achieve self-learning of hovering control for aerial robots,and compared with QAC and DDPG,it has higher learning efficiency,faster convergence speed,and smoother hovering effect.
卓浩泽;杨忠;吴吉莹;何加辉
南京航空航天大学自动化学院,江苏南京 211106||广西电网有限责任公司电力科学研究院广西电力装备智能控制与运维重点实验室,广西南宁 530000南京航空航天大学自动化学院,江苏南京 211106南京航空航天大学自动化学院,江苏南京 211106南京航空航天大学自动化学院,江苏南京 211106
空中机器人策略梯度算法动态监督器悬停控制强化学习
aerial robotpolicy gradient algorithmdynamic monitorhovering controlreinforcement learning
《控制理论与应用》 2026 (8)
1649-1657,9
广西电网有限责任公司科技项目(GXKJXM20230169),国家自然科学基金面上项目(61473144)资助.Supported by the Science and Technology Project of Guangxi Power Grid Co.,Ltd.(GXKJXM20230169)and the General Program of National Natural Science Foundation of China(61473144).
评论