LogoCLIP:基于零样本学习的商标外观异常检测模型OA
LogoCLIP:A Zero-Shot Method for Trademark Appearance Anomaly Detection
商标作为产品的重要标识,其外观质量对购物体验和品牌形象等具有深远影响.针对现有工业异常检测方法需使用大量数据训练模型,难以应用于数据量少、种类繁多、缺陷模式不固定的商标外观异常检测的问题,提出一种基于零样本学习的商标外观异常检测模型 LogoCLIP.该模型在视觉-语言预训练模型 CLIP 的基础上,引入可学习的对象无关文本提示,引导模型关注异常区域而非对象语义,以实现对不同对象通用的外观异常检测.提出一种基于注意力机制的跨模态特征融合模块,实现图像-文本特征对齐;通过设计门控融合适配器解决预训练数据与目标数据的分布偏移;借助辅助数据集,模型可在无需商标数据训练样本的情况下实现商标外观异常检测.在自建商标外观缺陷数据集上的实验结果表明,LogoCLIP 的图像级 AUROC 达到 92.9%,像素级AUROC 达到 99.3%;在 VisA、SDD、BTAD 3 个公开工业数据集上,模型同样取得了具有竞争力的检测结果,验证了其较强的泛化能力.
As an important identifier of products,the visual quality of trademarks has a profound impact on consumers'shopping experiences as well as brand reputation.However,existing appearance defect detection methods in the industry primarily rely on training models with large amounts of data,making it difficult to apply them to the field of trademark appearance anomaly detection,which involves a wide variety of products,limited data,and non-fixed defect patterns.To address this issue,this paper proposes LogoCLIP:A zero-shot learning-based model for trademark appearance anomaly detection.Based on the vision-language pre-training model CLIP,LogoCLIP introduces learnable object-agnostic text prompt to guide the model to focus on anomalous regions rather than object semantics,enabling generalized anomaly detection across different types of objects.Then,a cross-modal feature fusion method based on attention mechanisms is proposed to achieve image-text feature alignment,and a gated fusion adapter is employed to address the distribution shift between pre-trained data and target data.By utilizing auxiliary datasets,the model achieves anomaly detection for trademark appearances without requiring training samples from trademark data.Finally,Experimental results on trademark appearance defect dataset demonstrate that LogoCLIP achieves image-level AUROC and pixel-level AUROC scores of 92.9%and 99.3%respectively.The model also achieved competitive detection results on three public industrial datasets—VisA,SDD and BTAD,which verified its strong generalization ability.
施凯悦;任志刚
广东工业大学自动化学院,广东 广州 510006广东工业大学自动化学院,广东 广州 510006||广东省智能系统与优化集成重点实验室,广东 广州 510006
信息技术与安全科学
商标外观异常检测零样本跨模态特征融合
trademark appearanceanomaly detectionzero-shotcross-modalfeature fusion
《自动化与信息工程》 2026 (3)
32-41,10
广东省基础与应用基础研究基金项目(2024A1515011768)国家自然科学基金项目(62073088).
评论