面向大机动目标的高效强化学习拦截制导律研究OA
Research on Efficient Reinforcement Learning Interception Guidance Law for Large Maneuvering Targets
针对大机动目标拦截困难下目标加速度信息难以获取、传统制导律对高动态环境的适应性不足、强化学习算法在广域时空下的探索效率低及训练稳定性差的问题,本文提出了一种双随机网络蒸馏-真实近端策略优化算法的末制导律.首先针对三维末制导场景设计了马尔可夫决策过程,然后基于深度强化学习算法对高速飞行器进行训练.为了提高飞行器在广域时空下的探索效率,提出了双随机网络蒸馏,通过引入内部奖励激励高速飞行器探索未知状态.为了加快训练速度、保证训练的稳定性,利用真实近端策略优化算法的信赖域回滚目标函数替代裁剪目标函数,有效提高了算法的收敛速度,增强了训练的稳定性.仿真结果表明,面对大机动目标,本文所设计的末制导律具备泛化性和鲁棒性,且无需目标加速度信息即可实现高效拦截.
In response to the challenges posed by intercepting highly maneuverable targets,including difficulties in obtaining target acceleration information,insufficient adaptability of traditional guidance laws to high-dynamic environments,low exploration efficiency of reinforcement learning algo-rithms in wide spatial-temporal contexts,and poor training stability,this paper proposes a terminal guidance law based on a double random distillation network and real proximal policy optimization algo-rithm.Firstly,a Markov decision process is designed for the three-dimensional terminal guidance sce-nario.Subsequently,deep reinforcement learning algorithms are employed to train high-speed vehicles.To enhance the exploration efficiency of these vehicles within wide spatial-temporal con-texts,a double random distillation network is introduced that incentivizes the exploration of unknown states through internal rewards.To accelerate training speed and ensure training stability,the trust region rollback objective function of real proximal policy optimization algorithm is used to replace the clipping objective function for effectively improving the convergence rate while enhancing training sta-bility.Simulation results demonstrate that the proposed terminal guidance law exhibits generalization and robustness when faced with highly maneuverable targets,and can achieve efficient interception without requiring target acceleration information.
张烨;涂远刚;郭正玉;王靖宇
西北工业大学 航天学院,西安 710072||空基信息感知与融合全国重点实验室,河南 洛阳 471009西北工业大学 航天学院,西安 710072中国空空导弹研究院,河南 洛阳 471009||空基信息感知与融合全国重点实验室,河南 洛阳 471009西北工业大学 航天学院,西安 710072||空基信息感知与融合全国重点实验室,河南 洛阳 471009
军事科技
末制导律强化学习训练效率近端策略优化随机网络蒸馏飞行器
terminal guidance lawreinforcement learningtraining efficiencyproximal policy optimizationrandom distillation networkvehicle
《航空兵器》 2026 (2)
44-53,10
国家自然科学基金项目(02564897)
评论