基于人工势场法的无人机集群突防博弈研究OA
Research on Penetration Games of UAV Swarms Based on the Artificial Potential Field Method
针对三维空域内红蓝四旋翼集群多对多零和突防博弈中高维连续决策难以收敛、探索效率低下以及策略鲁棒性欠佳的问题,本文提出一种融合人工势场先验与对手策略预测的多智能体深度确定性策略梯度算法(MADDPG)求解框架.首先,将无人机三自由度运动学嵌入完全信息微分博弈,设计"任务-威胁-协同"三阶收益,并引入势场可微势能,把稀疏终端奖励转化为稠密梯度信号,实现"趋利避害"先验的显式表征;其次,构建势场引导的混合探索,在线以势能方向调制奥恩斯坦-乌伦贝克过程(OU)噪声,离线以势场正则平滑目标Q值,提升样本利用率并抑制过估计;最后,集成轻量化对手策略预测器,在Actor梯度中引入元博弈项,使红方策略更新时即最小化对手预期收益,主动破坏敌方决策一致性,加速逼近纳什均衡.仿真结果表明,本文所提方法在2V2与4V4密集对抗中胜率稳定超过90%,系统诱导蓝方产生冗余加速度与能量耗散,持续撕开时空间隔完成无碰撞突防,显著优于无预测的MADDPG,验证了框架在多对多零和博弈中的可扩展性、实时性与鲁棒性.
Targeting the challenges of high-dimensional continuous decision non-convergence,low exploration efficiency,and insufficient policy robustness in multi-to-multi red-blue quadrotor swarm zero-sum penetration games within three-dimensional airspace,this paper proposes a Multi-Agent Deep Deterministic Policy Gradient(MADDPG)solution framework integrating artificial potential field priors with opponent strategy prediction.First,the three-degree-of-freedom UAV kinematics are embedded into a complete-information differential game,designing a"mission-threat-cooperation"three-tier reward structure,and introducing differentiable potential field energy to transform sparse terminal rewards into dense gradient signals,achieving explicit representation of the"seeking-advantage-avoiding-disadvantage"prior.Second,a potential field-guided hybrid exploration mechanism is constructed,online modulating Ornstein-Uhlenbeck process(OU)noise using potential energy directions,and offline smoothing target Q-values with potential field regularization,improving sample utilization and suppressing overestimation.Furthermore,a lightweight opponent strategy predictor is integrated,introducing a meta-game term into the Actor gradient,enabling red-team policy updates to simultaneously minimize opponent expected payoffs,proactively disrupting enemy decision consistency and accelerating convergence to Nash equilibrium.Simulation results demonstrate that the proposed method achieves stable win rates exceeding 90%in 2v2 and 4v4 dense confrontations,systematically induces blue team to generate redundant accelerations and energy dissipation,continuously creates spatial-temporal gaps to complete collision-free penetration,significantly outperforming MADDPG without prediction,validating the framework's scalability,real-time performance,and robustness in multi-to-multi zero-sum games.
王瑞昌;石琛;张科;呼卫军;马先龙
西北工业大学 航天学院,陕西 西安 710072上海机电工程研究所,上海 201109西北工业大学 航天学院,陕西 西安 710072西北工业大学 航天学院,陕西 西安 710072西北工业大学 航天学院,陕西 西安 710072
航空航天
人工势场多对多零和突防博弈强化学习策略预测
artificial potential fieldmany-versus-many zero-sum penetration gamereinforcement learningpolicy prediction
《空天防御》 2026 (2)
8-17,10
中国航天科技集团有限公司上海航天科技创新基金项目(SAST2022-006)
评论