首页|期刊导航|测试科学与仪器|基于自适应原型与语义感知的长尾表征学习算法

基于自适应原型与语义感知的长尾表征学习算法OA

Long-tailed representation learning algorithm based on adaptive prototypes and semantic awareness

中文摘要英文摘要

标注稀缺与长尾分布不均衡是工业设备监测领域所面临的关键技术挑战.现有自监督学习方法在复杂工况下易受样本数量偏差与语义混淆影响,导致对稀疏关键状态的表征能力受限.本文提出一种协同进化的原型对比学习框架(Co-evolutionary prototypical contrastive learning,EPCL),通过从粗粒度语义发现到细粒度判别增强的递进式学习,实现对长尾数据内在结构的深度解析.该框架设计了基于最优传输理论的自适应原型聚类算法,以数据驱动的动态先验实现无偏建模.而后提出语义感知的负样本层次化调控机制,采用原型一致性约束与自适应加权策略,实现判别边界优化的同时缓解类别不均衡.在多个公开长尾视觉基准(CIFAR10-LT、CIFAR100-LT、ImageNet-100-LT)以及工业故障诊断数据集CWRU上的实验结果表明,EPCL相较于SimCLR、SwAV等15种主流自监督方法,在线性评估与少样本分类任务中均取得显著提升,其在CIFAR100-LT上尾部类别准确率比SimCLR提高4.56个百分点.消融实验与可视化结果验证了所提机制的有效性与泛化能力.本研究为无监督条件下长尾测量数据的表征学习提供了具有潜力的新思路及实施方法.

Label scarcity and long-tailed distribution imbalance are significant challenges in industrial equipment monitoring.Currently,self-supervised learning methods are affected by sample quantity bias and semantic confusion under complex operating conditions,which limits their ability to represent sparse critical states.To address these issues,we propose a co-evolutionary prototypical contrastive learning(EPCL)framework.Through progressive learning from coarse-grained semantic discovery to fine-grained discriminative enhancement,this framework enables an in-depth analysis of the intrinsic structure of long-tailed data.Specifically,an adaptive prototype-based clustering algorithm based on optimal transport theory is introduced,thereby achieving unbiased representation learning through data-driven dynamic priors.Furthermore,a semantic-aware and hierarchical negative sample weighting scheme is designed to optimize discriminative boundaries while mitigating class imbalance by enforcing prototype consistency constraints and employing an adaptive weighting strategy.Extensive experiments were conducted on several public long-tailed visual benchmarks,including CIFAR10-LT,CIFAR100-LT,and ImageNet-100-LT,as well as the industrial fault diagnosis dataset.The results demonstrated that the EPCL achieved better performance than fifteen mainstream self-supervised methods(e.g.,SimCLR and SwAV)in both linear evaluation and few-shot classification tasks.On the CIFAR100-LT dataset,the EPCL improved the tail-class accuracy by 4.56%compared to SimCLR.Ablation studies and visualization results verified the effectiveness and generalization ability of the framework.This work offers a promising insight and practical solution for representation learning from unlabeled long-tailed measurement data.

李甜甜;薛震;张亮亮;连旭

中北大学 数学学院,山西 太原 030051中北大学 数学学院,山西 太原 030051中北大学 数学学院,山西 太原 030051中北大学 数学学院,山西 太原 030051

长尾分布自监督学习对比学习语义感知自适应原型聚类表征学习

long-tailed distributionself-supervised learningcontrastive learningsemantic awarenessadaptive prototype clusteringrepresentation learning

《测试科学与仪器》 2026 (2)

331-343,13

The work was supported by the National Natural Science Foundation of China(No.12401703)and the Fundamental Research Program of Shanxi Province(Nos.202203021211088,202403021221109,202403021212256).

10.62756/jmsi.1674-8042.2026028

评论