无人机视觉语言导航模型综述:从感知理解到智能决策OA
A Review of Vision-Language Navigation Models for UAVs:From Perceptual Comprehension to Intelligent Decision-Making
多模态大语言模型的兴起为视觉-语言-导航范式奠定了基础,它将视觉感知、自然语言理解与导航控制纳入同一策略框架.无人机领域迅速借鉴这一思路,尝试使飞行器能够理解自然语言指令、在三维场景中推理并做出飞行决策.相比传统分模块导航,基于多模态大语言模型的端到端框架能够同步处理语言与视觉信号,将感知信息直接映射为控制指令.然而,无人机视觉-语言-导航的研究目前缺乏系统梳理.该文对其最新进展做了综合回顾:从早期模块化方案到以推理为核心的视觉-语言-行动模型,剖析视觉、语言、控制信息如何逐步融合以提升自主导航能力;随后梳理现有数据集与评测标准,涵盖室内、室外复杂场景的仿真任务和真实飞行轨迹,指标包括成功率、耗时及语义理解能力;最后归纳关键挑战,包括跨模态对齐困难、动态环境下的实时响应能力不足、标注成本高,以及复杂场景下决策鲁棒性差.该文呈现了无人机自主导航研究的新路线与未来的研究方向,指出大规模多模态大语言模型提升无人机智能决策与可解释性方面的前景,可为无人机安全高效自主飞行的研究与实践提供参考.
The emergence of multimodal large language models(MLLMs)has laid the foundation for the vision-language-navigation paradigm,which integrates visual perception,natural language understanding,and navigation control within a unified strategic framework.This paradigm has been rapidly adopted in the UAV domain,attemp-ting to enable UAVs to understand natural language instructions,reason in three-dimensional environments,and make flight decisions.Compared with traditional modular navigation approaches,the end-to-end framework based on MLLMs can simultaneously process linguistic and visual signals,directly mapping perceptual information into control commands.However,a systematic review of UAV-VLN remains scarce.This paper presents a comprehen-sive review of recent advances in this area:from early modular solutions to reason-centric vision-language-action models.It elucidates how visual,linguistic,and control information are progressively integrated to enhance autono-mous navigation capabilities.It further summarizes existing datasets and evaluation protocols,including simulation tasks in both indoor and outdoor complex environments as well as real-world UAV flight trajectories,with evaluation metrics covering success rate,time cost,and semantic comprehension.Finally,it identifies key challenges,inclu-ding difficulties in cross-modal alignment,insufficient real-time responsiveness in dynamic environments,high an-notation costs,and poor decision-making robustness in complex scenarios.This paper outlines new pathways and future research directions for UAV autonomous navigation research.It highlights the potential of MLLMs in enhan-cing intelligent decision-making and interpretability of UAVs,and serves as a reference for research and practice in achieving safe and efficient autonomous flight of UAVs.
王子豫;杜宸旭;刘洋
香港科技大学(广州)智能交通学域,广东 广州 511453西南交通大学 交通运输与物流学院,四川 成都 611756清华大学 车辆与运载学院,北京 100084
信息技术与安全科学
无人机视觉-语言导航多模态大语言模型智能决策
UAVvision-language navigationmultimodal large language modelsintelligent decision-making
《华南理工大学学报(自然科学版)》 2026 (6)
173-182,10
北京市昌平区"U35培苗资助计划"项目(CHPU35202504003)国家自然科学基金项目(52572334,T2588101,52221005,52220105001)清华大学智能绿色车辆与交通全国重点实验室自主研究课题(ZZ-GG-20250403)Supported by the"U35 Pei Miao Funding Program"of Changping District,Beijing(CHPU35202504003),the National Natural Science Foundation of China(52572334,T2588101,52221005,and 52220105001)and the Independent Research Project of the State Key Laboratory of Intelligent Green Vehicle and Mobility,Tsinghua University(ZZ-GG-20250403)
评论