融合卷积与自注意力机制的可见光-红外目标检测OA
Visible-Infrared Object Detection Integrating Convolution and Self-attention Mechanism
为提升红外与可见光目标检测方法在光照不均、目标重叠、目标尺度多样等复杂环境下的检测性能,将卷积神经网络与视觉注意力网络相结合,设计了一种基于卷积与自注意力机制的多模态目标检测模型.模型针对红外和可见光图像特点,采用轻量级网络设计理念构建多分支特征提取结构.可见光分支以分组卷积单元为核心从局部视角逐层提取图像中的纹理细节特征,红外分支以基于窗口的高效率多头自注意力模块为基础从全局视角提取图像中的显著特征.针对不同模态特征,设计跨模态信息融合结构,通过关键特征增强以及动态调整融合权重的方式实现红外和可见光特征互补.同时,为提升模型的尺度不变性,引入多尺度特征过滤融合结构,分别从通道和空间位置角度缓解不同维度特征间信息冲突.在数据集上的实验结果表明,与当前同类方法相比,该方法在复杂场景中检测精度达到85.3%,且具备更高的泛化能力和鲁棒性.
In order to improve the detection performance of infrared and visible light target detection methods in complex environments such as uneven lighting,overlapping objects,and diverse object scales,a multimodal object detection model based on convolution and self-attention mechanisms is designed by combining convolutional neural networks with visual attention networks.The model adopts a lightweight network design concept to construct a multi-branch feature extraction structure based on the characteristics of infrared and visible light images.The visible light branch uses group convolution units as the core to progressively extract image texture details feature from a local perspective and the infrared branch utilizes an efficient multi-head self-attention module based on windows to extract significant image features from a global perspective.For different modal features,a cross-modal information fusion structure is designed to achieve complementarity between infrared and visible light features through key feature enhancement and dynamic adjustment of fusion weights.Meanwhile,to enhance the scale invariance of the model,a multi-scale feature filtering fusion structure is introduced to alleviate information conflicts between features of different dimensions from both channel and spatial location perspectives.Experiments on standard datasets show that compared with current similar methods,this method achives the detection accuracy of 85.3%in complex scenes,and it also has higher generalization ability and robustness.
周聪敏;吴斌鹏
郑州财经学院 信息工程学院,郑州 450049郑州大学 计算机与人工智能学院,郑州 450001
信息技术与安全科学
可见光-红外目标检测卷积网络视觉注意力网络特征融合
visible-infrared object detectionconvolutional networkvisual attention networkfeature fusion
《电讯技术》 2026 (8)
1308-1316,9
国家自然科学基金资助项目(51977074)河南省社科联调研课题(SKL2019888)
评论