首页|期刊导航|实验技术与管理|融合图神经网络与课程学习的DRL多无人机路径规划实验

融合图神经网络与课程学习的DRL多无人机路径规划实验OA

Experiment of deep reinforcement learning for multi-UAV path planning with graph neural networks and curriculum learning

中文摘要英文摘要

针对有限空域内多无人机在三维密集障碍环境中的路径规划与避障难题,以及传统全局规划动态变化适应性不足等问题,提出了一种基于图神经网络边特征耦合编码与三阶段渐进式课程学习的多智能体协同决策框架.在该框架下,通过深度建模无人机与障碍物间的边关系特征实现精准局部感知,并采用三阶段渐进式训练策略提升模型收敛速度及其鲁棒性.仿真实验结果表明,该框架可显著缩短平均路径长度、提升避障成功率及提高任务完成效率,同时为学生开展多智能体协同决策的系统化研究提供了有效的实践验证与教学支持.

[Objective]With the widespread application of unmanned aerial vehicles(UAVs)in disaster rescue,industrial inspection,and other scenarios,multi-UAV path planning in constrained airspace faces dual challenges in avoiding dense static obstacles and improving operational efficiency.Traditional path planning methods based on environmental priors struggle to adapt to dynamically generated scenarios with randomly distributed obstacles.Existing reinforcement learning algorithms predominantly rely on simplified two-dimensional planar assumptions,neglecting three-dimensional(3D)spatial constraints for obstacle avoidance.To address these limitations,this study proposes a collaborative decision-making framework for multi-UAV path planning in 3D static dense obstacle environments by integrating graph neural network(GNN)architecture optimization with progressive curriculum learning(CL).[Methods]First,a 3D path planning model is formulated based on the Markov decision process by incorporating altitude dimensions into a state representation and designing a node-type identification mechanism.This design enables UAVs to distinguish heterogeneous characteristics between themselves and surrounding obstacles.To address the limitations of conventional GNNs in spatial relationship modeling,this study couples edge features,including relative velocity,position,and distance,with neighbor node features,such as relative centroid position,velocity,and type identifiers.These features are then fused using multilayer perceptrons to generate joint representations.This approach replaces the linear superposition of independently encoded features commonly used in existing algorithms,thereby enhancing the network's capability to analyze complex spatial distributions of obstacles.Second,a reward function that balances safety and efficiency is formulated by integrating multidimensional metrics,including target proximity,first-arrival time,dwell duration,velocity alignment,collision risk,and proximity penalties.This design guides UAVs to achieve optimal trade-offs between obstacle avoidance and navigation objectives,thereby improving trajectory rationality and policy convergence speed.Third,a three-stage progressive training framework is developed,transitioning from sparse to dense obstacle scenarios.UAVs initially learn basic obstacle avoidance strategies in simplified environments,then gradually progress to moderate-difficulty environments,and ultimately generate cooperative paths balancing safety and efficiency in complex obstacle configurations.This methodology addresses suboptimal policy issues caused by excessive exploration in high-dimensional environments.Finally,a 3D multi-UAV path planning test environment is established using the PyBullet high-fidelity physics simulation platform,featuring randomly distributed static obstacles with varying density levels.[Results]Experimental results demonstrate that the proposed edge-couple informative multi-agent proximal policy optimization(EC-InforMAPPO)framework outperforms baseline algorithms across all difficulty levels.Its edge feature encoding mechanism,which couples relative motion parameters and spatial relationships,enhances trajectory safety in dense obstacle environments,offering a novel technical pathway for environmental perception modeling in multi-agent systems.Additionally,the progressive curriculum learning framework enhances policy stability in challenging scenarios.The EC-InforMAPPO-CL framework achieves higher obstacle avoidance success rates and faster convergence than direct training under equivalent computational resources.This establishes a reusable training paradigm for reinforcement learning in high-dimensional state spaces.[Conclusions]This study proposes a collaborative decision-making framework that combines edge feature coupling based on GNNs with progressive CL to address challenges in multi-UAV path planning in three-dimensional dense obstacle environments.The findings provide new insights and technical support for intelligent collaborative navigation of multiple UAVs in complex environments,holding significant application potential and practical value.

傅明建;陈文涛;卓晓鑫;陈恒升;陈飞

福州大学 计算机与大数据学院,福建 福州 350108||福州大学 网络信息安全与计算机技术国家级实验教学示范中心,福建 福州 350108福州大学 计算机与大数据学院,福建 福州 350108福州大学 计算机与大数据学院,福建 福州 350108福州大学 计算机与大数据学院,福建 福州 350108福州大学 计算机与大数据学院,福建 福州 350108||福州大学 网络信息安全与计算机技术国家级实验教学示范中心,福建 福州 350108

信息技术与安全科学

无人机多智能体强化学习图神经网络路径规划课程学习

UAVmulti-agent reinforcement learninggraph neural networkpath planningcurriculum learning

《实验技术与管理》 2026 (5)

136-144,9

福建省本科高校教育教学研究项目(FBJY20250314,FBJY20250175,FBJY20250093)福建省自然科学基金面上项目(2026J001253)福州大学研究生教育教学改革项目(FYAI2024010)

10.16791/j.cnki.sjg.2026.05.017

评论