低空视角下改进无人机小目标检测算法OA
Research on Improved Small-Object Detection Algorithm for UAVs from Low-Altitude Perspectives
针对低空经济快速发展背景下,无人机在农业、物流及应急救援等领域对高精度、轻量化目标检测的迫切需求,针对低空视角下目标像素占比低、易受环境遮挡及透视形变严重等挑战,基于RT-DETR(Real-Time Detection Transformer)算法,提出一种改进的跨尺度对齐与位置编码增强检测模型(CAPE-RT-DETR).首先,为克服传统静态卷积核在复杂背景下特征提取灵活性不足的问题,设计了融合动态卷积核生成与门控特征选择的特征增强模块(C2ML),通过利用大核预测器动态生成空间自适应的卷积核,并结合门控特征选择机制剔除冗余背景信息,显著强化了模型对关键特征的提取与筛选能力;其次,针对航拍视角下的几何畸变与空间非均匀形变,将可学习位置编码与多头自注意力机制相结合,构建了增强型位置感知交互模块(AIFP),通过端到端学习空间先验信息,有效提升了模型对低空特异性空间结构的感知灵敏度与定位精度;最后,针对多尺度特征融合中因简单上采样导致的像素位移难题,引入跨尺度特征校准模块(CSFC),利用金字塔场景解析结构整合稀疏全局上下文,并辅以双路卷积与网格采样机制显式补偿跨尺度对齐偏差,实现了语义信息的一致性表达.在ALU与VisDrone2019数据集上的实验结果表明,CAPE-RT-DETR在参数量、精度与模型大小等方面均优于基线算法,同时消融实验验证了3个改进模块的有效性与协同性.该研究为复杂场景下无人机实时目标检测提供一种高精度、轻量化的方法基础与理论支持.
In the context of the rapidly growing low-altitude economyand the urgent need for high-precision,lightweight UAV object detection in agriculture,logistics,and emergency rescue,this paper tackles the challenges of low target pixel occupancy,environmental occlusion,and severe perspective distortion in low-altitude imagery.Based on the RT-DETR algorithm,an enhanced detection model—Cross-scale Alignment and Position Encoding Enhanced RT-DETR(CAPE-RT-DETR)—is proposed.Firstly,to overcome the limitation of traditional static convolution kernels in terms of feature extraction flexibility under complex backgrounds,this paper proposes a feature enhancement moduleintegrating dynamic convolution kernel generation and gated feature selection,termed C2ML.By utilizing a Large Kernel Predictor(LKP)to dynamically generate spatially adaptive convolution kernels,and combining it with a gated feature selection mechanism to eliminate redundant background information,the module significantly enhances the model's ability to extract and filter critical features.Secondly,to address the geometric distortion and spatially non-uniform deformation characteristic of aerial perspectives,a learnable positional encoding is integrated with the multi-head self-attention mechanism to construct an enhanced position-aware interaction module,termed AIFP.By learning spatial prior information in an end-to-end manner,this module effectively improves the model's perceptual sensitivity and localization accuracy with respect to the low-altitude-specific spatial structures.Finally,to resolve the pixel misalignment problem caused by simple upsampling in multi-scale feature fusion,a a cross-scale feature calibration(CSFC)module is introduced.This module utilizes a pyramid scene parsing structure to integrate sparse global context and employs a dual-path convolution and grid sampling mechanism to explicitly compensate for cross-scale alignment biases,thereby achieving consistent representation of semantic information.Experimental results on the ALU and VisDrone2019 datasets demonstrate that CAPE-RT-DETR outperforms the baseline algorithm in terms of parameter count,accuracy,and model size.Meanwhile,ablation experiments validate the effectiveness and synergy of the three improved modules.This research provides a high-precision and lightweight methodological foundation and theoretical support for real-time UAV object detection in complex scenarios.
张杰;董春彤;裴玉龙;何庆龄
宁德师范学院 信息工程学院,福建 宁德 352100||东北林业大学 土木与交通学院,黑龙江 哈尔滨 150040东北林业大学 土木与交通学院,黑龙江 哈尔滨 150040东北林业大学 土木与交通学院,黑龙江 哈尔滨 150040兰州交通大学 交通运输学院,甘肃 兰州 730070
交通工程
低空交通目标检测无人机RT-DETR
low-altitude trafficobject detectionunmanned aerial vehicleRT-DETR
《华南理工大学学报(自然科学版)》 2026 (6)
193-204,12
国家自然科学基金重点项目(51638004)福建省自然科学基金项目(2023J011093)东北林业大学中央高校基本科研业务费专项资金项目(2572023CT21-02).Supported by the the Key Project of National Natural Science Foundation of China(51638004)and the Natural Science Foundation of Fujian Province(2023J011093)
评论