多模态蒸馏驱动的单RGB鲁棒目标检测方法OA
Multi-modal Distillation-driven Robust Object Detection Method with Single RGB
针对复杂光照、低可见度或局部遮挡条件下,单RGB轻量化目标检测模型鲁棒性不足,以及多模态传感器方案在边缘部署中存在硬件成本高、标定复杂和功耗受限等问题,提出一种多模态蒸馏驱动的单RGB鲁棒目标检测方法.该方法构建了训练阶段由多模态教师网络引导、推理阶段仅保留单RGB学生模型的异构教师-学生蒸馏框架:训练阶段以RGB图像及辅助模态图像构建教师网络,通过退化建模、跨模态语义投影、跨模态一致性约束和模态补偿注意力机制,将热红外、深度等辅助模态中的结构信息,迁移至仅接收RGB输入的轻量学生网络;推理阶段仅保留单RGB学生模型,以实现低成本部署.为增强实验可验证性,围绕自建的多模态训练数据集USAB-180K、公开评测基准KAIST数据集进行对比、缺失模态泛化、消融和部署等实验.实验结果表明,本文方法在KAIST测试集上的Macro-F1较单RGB轻量基线模型提高了3.4个百分点,误检率下降了5.6个百分点,mAP@50由62.4%提升至72.1%;在1080p,30 f/s视频流输入条件下,模型部署于RV1106工业视觉平台时,其平均端到端时延为31.8 ms,推理功耗为206 mW.该方法为资源受限边缘终端提升复杂环境下的目标检测性能稳定性提供了一种可行思路.
To address the insufficient robustness of lightweight single-RGB object detection models under complex illumination,low visibility,or partial occlusion conditions,as well as the challenges of high hardware cost,complex calibration,and limited power consumption faced by multimodal sensor solutions in edge deployment,this paper proposes multi-modal distillation-driven robust object detection method with single RGB.The method constructs a heterogeneous teacher-student distillation framework in which a multimodal teacher network guides the training phase while only the single-RGB student model is retained for inference.During training,the teacher network is built with RGB images and auxiliary modal images,transferring structural information from thermal infrared,depth,and other auxiliary modalities to a lightweight student network that receives only RGB inputs through degradation modeling,cross-modal semantic projection,cross-modal consistency constraints,and modal-compensation attention mechanisms.During inference,only the single-RGB student network is retained to enable low-cost deployment.To enhance experimental verifiability,comparative experiments,missing-modality generalization experiments,ablation studies,and deployment experiments are conducted on the self-built multimodal training dataset USAB-180K and the public benchmark KAIST dataset.Experimental results demonstrate that the proposed method achieves a 3.4 percentage point improvement in Macro-F 1 over the single-RGB baseline model on the KAIST test set,reduces the false positive rate by 5.6 percentage points,and increases mAP@50 from 62.4%to 72.1%.Under 1080p,30 f/s video stream input,when deployed on the RV1106 industrial vision platform,the model achieves an average end-to-end latency of 31.8 ms and an inference power consumption of 206 mW.The proposed method provides a feasible approach for enhancing the stability of object detection performance on resource-constrained edge devices in complex environments.
艾炎;叶廷东
广东轻工职业技术大学人工智能学院,广东 广州 510300广东轻工职业技术大学人工智能学院,广东 广州 510300
信息技术与安全科学
目标检测多模态蒸馏单RGB视觉知识蒸馏边缘部署
object detectionmulti-modal distillationsingle-RGB visionknowledge distillationedge deployment
《自动化与信息工程》 2026 (4)
8-16,9
广东省普通高校重点科研平台和项目(2023ZDZX1053)
评论