首页|期刊导航|自动化与信息工程|多尺度自监督特征融合的工业视觉缺陷检测方法

多尺度自监督特征融合的工业视觉缺陷检测方法OA

Multi-scale Self-supervised Feature Fusion Method for Industrial Visual Defect Detection

中文摘要英文摘要

针对产品品类切换后视觉缺陷检测模型需重新训练与调整参数的问题,提出一种多尺度自监督特征融合的工业视觉缺陷检测方法.该方法以冻结的自监督视觉主干网络通过多Hook并行抓取浅、中、深3层patch token特征,分层构建核心子集记忆库;通过等权得分级融合机制将各层最近邻距离合成统一的异常图像,从而避免模型逐类调整参数与重新训练.在MVTec-AD数据集全部15类、3个随机种子下,图像级AUROC达0.987 4±0.0017,跨种子标准差仅为同主干基线的11%~20%,区域级PRO达0.8163±0.012 1.归因消融实验结果表明:跨种子标准差减小的主要原因是多层Hook的等权得分级融合机制,全部15类上多层相对单层L9的标准差降幅约为45%.零样本迁移至VisA数据集全部12类时,P-AUROC达0.992 9;Real-IAD数据集前6类子集的P--AUROC达0.9847.在自建的发动机缸套缺陷图像数据集上,P-AUROC达0.967 3±0.001 5.在半精度模式下单图推理耗时约为25 ms,可满足一般中速生产线的节拍要求.

Aiming at the problem that visual defect detection models require retraining and parameter tuning after product category switching,we propose an industrial visual defect detection method based on multi-scale self-supervised feature fusion.The method employs a frozen self-supervised visual backbone network to extract shallow,middle,and deep patch token features in parallel via multiple hooks,and hierarchically constructs a core subset memory bank for each layer.It then fuses the nearest neighbor distances from all layers into a unified anomaly map using an equal-weight score-level fusion mechanism,thereby avoiding category-wise parameter tuning and retraining.On all 15 categories of the MVTec-AD dataset under 3 random seeds,the image-level AUROC reaches 0.987 4±0.001 7,with a cross-seed standard deviation that is only 11%-20%of that of the same backbone baseline,while the region-level PRO achieves 0.816 3±0.012 1.Ablation study results attribute the reduction in cross-seed standard deviation primarily to the equal-weight score-level fusion mechanism of multiple hooks;the standard deviation reduction of multi-layer compared to single-layer L9 across all 15 categories is approximately 45%.When zero-shot transferred to all 12 categories of the VisA dataset,the P-AUROC reaches 0.992 9;on the first 6 categories of the Real-IAD dataset,the P-AUROC achieves 0.984 7.On a self-collected engine cylinder liner defect image dataset,the P-AUROC reaches 0.967 3±0.001 5.In half-precision mode,the inference time per image is about 25 ms,which can meet the takt time requirements of typical medium-speed production lines.

陈永彬;黄嘉思;黄国健

广东机电职业技术学院,广东 广州 510550广东机电职业技术学院,广东 广州 510550广东机电职业技术学院,广东 广州 510550

信息技术与安全科学

工业视觉缺陷检测自监督学习多尺度特征零样本迁移

industrial visualanomaly detectionself-supervised learningmulti-scale featureszero-shot transfer

《自动化与信息工程》 2026 (4)

17-25,9

广州市青年科技人才托举工程项目(QT-2025-010)

10.12475/aie.20260403

评论