基于解耦时空特征融合的小目标无人机检测OA
Decoupled spatiotemporal feature fusion based small UAV object detection
针对视频检测远距离、小尺寸无人机目标时易出现漏检和误检的问题,提出一种解耦时空特征融合检测网络(DS2F-DN).该网络以YOWOv2为基础,通过差异化分辨率预处理,将高清关键帧与低分辨率帧序列分别输入二维和三维卷积分支进行处理,可以综合利用图像的空间信息和视频帧序列的时序信息.具体地,在二维卷积分支中设计了基于小波变换的特征解耦模块(WTDM),用于分解空间特征包含的高低频信息.在三维卷积分支中,使用反卷积实现时空特征的多尺度化,并设计了多尺度特征交互机制,用于提升时空特征表达能力.进一步,针对WTDM输出的二维空间解耦特征与三维卷积分支输出的时空压缩特征,在通道编码器中采用并行逐点卷积结构,实现特征的降维和融合,同时引入缩放点积自注意力机制,增强模型的时空建模能力.实验结果表明,所提DS2F-DN在无人机视频数据集Drone-detection-dataset上能够以23.4 ms的单帧推理时延(约42.7 FPS)实现约71.51%mAP的检测精度,综合性能优于现有其它方法,可实现小目标无人机高精度实时检测.
In response to the issues of missed and false detections when detecting small-sized drones at long distances in video,this paper proposes a Decoupled Spatiotemporal Feature Fusion Detection Network(DS2F-DN)which is built on the YOWOv2(You Only Watch Once version 2)framework and employs a differential resolution preprocessing,wherein high-definition key frames and low-resolution frame sequences are separately input into two-dimensional(2D)and three-dimensional(3D)convolutional branches for processing.This approach allows for the comprehensive utilization of spatial information from images and temporal information from video frame sequences.Specifically,we design a Wavelet Transform Decoupling Module(WTDM)within the 2D convolutional branch to decompose the high-and low-frequency information contained in the spatial features.In the 3D convolutional branch,we adopt transposed convolution to achieve multi-scaling of spatiotemporal features,and design a multi-scale feature interaction strategy to enhance the representation ability of spatiotemporal features.Furthermore,regarding the 2D spatial decoupled features output by WTDM and the spatiotemporally compressed features from the 3D convolutional branch,we imple-ment a parallel pointwise convolution structure in the channel encoder to achieve dimensionality reduction and feature fusion.Additionally,we introduce a scaled dot-product self-attention mechanism to improve the model's spatiotemporal modeling capabilities.Experimental results show that the DS2F-DN can achieve a detection accuracy of 71.51%mAP with a frame inference latency of 23.4 ms(approximately 42.7 FPS)on the drone video dataset,Drone-detection-dataset,which outperforms existing methods in overall performance and achieves high-precision real-time detection of small drone targets.
阳小兵;李嘉冰;李钊;丁汉清
西安电子科技大学 网络与信息安全学院,陕西 西安 710126西安电子科技大学 网络与信息安全学院,陕西 西安 710126西安电子科技大学 网络与信息安全学院,陕西 西安 710126郑州轻工业大学 电子信息学院,河南 郑州 450000
信息技术与安全科学
深度学习视频目标检测无人机检测时空特征融合
deep learningvideo object detectionUAV detectionspatiotemporal feature fusion
《西安电子科技大学学报(自然科学版)》 2026 (3)
19-31,13
河南省科技攻关项目(252102211120)国家自然科学基金(62072351,62202359,U23A20300)高等学校学科创新引智计划(B16037)
评论