基于自适应熵值的移动机器人无地图导航方法OA
Mapless navigation method for mobile robots based on adaptive-entropy
针对现有基于机器学习的无地图导航方法中所存在的机器人陷入局部最优后脱困能力差的问题,提出一种自适应熵值无地图导航方法.该方法首先根据机器人训练过程中的奖励波动来判定局部最优现象是否发生,然后在此基础上继续通过改进软演说评论家(SAC)算法中的熵值调节机制,达到在机器人陷入局部最优后自适应改变温度系数以强化其学习能力的效果,最终促使机器人成功学习到全局最优导航策略.仿真结果表明:相比现有方法,所提方法训练得到的导航策略具备更强的脱困能力和更高的导航成功率.
Aiming at the problem of poor ability to escape from local optima after a robot became trapped,which existed in existing mapless navigation methods based on machine learning,an adaptive entropy mapless navigation method was proposed.First,this method determined whether a local optimum phenomenon occurred according to the reward fluctuations during the robot's training process,and then on this basis,continued to improve the entropy adjustment mechanism in the soft actor-critic(SAC)algorithm to achieve the effect of adaptively changing the temperature coefficient to strengthen its learning ability after the robot became trapped in a local optimum,and finally promoted the robot to successfully learn a globally optimal navigation strategy.Simulation results show that compared with existing methods,the navigation strategy trained by the proposed method has stronger escape ability and higher navigation success rate.
熊体凡;李泓辰;王书亭;谢远龙;胡倚铭
华中科技大学机械科学与工程学院,湖北武汉 430074华中科技大学机械科学与工程学院,湖北武汉 430074华中科技大学机械科学与工程学院,湖北武汉 430074华中科技大学机械科学与工程学院,湖北武汉 430074华中科技大学机械科学与工程学院,湖北武汉 430074
信息技术与安全科学
移动机器人无地图导航深度强化学习局部最优SAC算法
mobile robotsmapless navigationdeep reinforcement learninglocal optimasoft actor-critic(SAC)algorithm
《华中科技大学学报(自然科学版)》 2026 (5)
98-104,7
国家自然科学基金资助项目(52275488).
评论