基于强化学习的无人攻击机与诱饵机协同航迹规划OA
Reinforcement Learning-Based Cooperative Trajectory Planning for Unmanned Combat Aerial Vehicles and Decoy UAVs
在现代战争中,无人机协同作战已成为提升战场效能的关键技术.其中,无人战斗机(UCAV)与诱饵机(Decoy UAV)的协同作战模式因其战术价值受到广泛关注.针对无人攻击机与诱饵机协同打击敌方重点目标的任务,提出了一种基于近端策略优化(proximal policy optimization,PPO)算法的协同航迹规划方法.构建融合动态威胁评估的马尔可夫决策过程(Markov decision process,MDP)模型,集成无人机运动学与战场环境约束,设计状态、动作空间及分层奖励函数.仿真实验表明,所提方法能有效引导攻击机与诱饵机在复杂战场环境中实现高效协同,显著提升任务成功率并降低被敌方防空系统拦截的风险,为无人机协同作战的智能化路径规划提供了理论与技术支撑.
Unmanned aerial vehicle(UAV)cooperative combat is crucial in modern warfare.The cooperative mode between unmanned combat aerial vehicles(UCAVs)and decoy UAVs has gained significant attention due to its tactical value.This paper proposes a cooperative trajectory planning method based on the proximal policy optimization(PPO)algorithm for UCAV and decoy UAV strike missions against key enemy targets.We construct a Markov decision process(MDP)model incorporating dynamic threat assessment,integrating UAV kinematics and battlefield constraints,and design the state/action spaces and a hierarchical reward function.Simulation results demonstrate that the proposed method effectively guides UCAVs and decoys to achieve efficient cooperation in complex environments,significantly increasing mission success rates while reducing interception risks from enemy air defense systems.This provides theoretical and technical support for intelligent path planning in UAV cooperative operations.
祁昊哲;郑明发;胡小荣;杨楠
空军工程大学 空管领航学院,陕西 西安 710051空军工程大学 基础部,陕西 西安 710051军事科学院 国防科技创新研究院,北京 100071空军工程大学 空管领航学院,陕西 西安 710051
航空航天
无人机协同作战航迹规划强化学习PPO算法马尔可夫决策
UAVs cooperative operationstrajectory planningreinforcement learningproximal policy optimization(PPO)
《现代防御技术》 2026 (3)
71-81,11
评论