对比学习优化的视频异常内容检测OA
Optimizing Video Anomaly Detection with Contrastive Learning
基于多示例学习的弱监督视频内容异常检测方法是当前异常检测的主流方法,该类方法以示例异常分数作为学习目标,研究优化示例异常分数或最高异常分数选择的问题.但这种学习目标弱化了视频包标签的重要性,而示例分数的学习则因缺乏标注信息只能通过无监督的方式进行学习,限制了异常检测的效果.为了解决上述问题,本文将视频包标签作为学习目标,提出一种端到端对比学习嵌入的视频异常检测网络.首先,该网络利用膨胀3D卷积网络(Inflated 3D ConvNet,I3D)进行特征提取并采用双向编码预训练模型BERT(Bidirectional Encoder Representations from Transform-ers)对特征进行优化,然后通过深度神经网络(Deep Neural Network,DNN)拟合示例分数,之后引入Noisy-or模型将示例分数通过符合多示例约束的方式投影为包标签并作为网络学习的最终目标,构造包特征到包标签的学习模型,进而通过二元交叉熵利用包的真实标签完成端到端的网络学习.其次,为了解决细粒度标注缺乏带来的示例分数学习欠佳的问题,本文进一步引入细粒度序列距离(Fine-grained Sequence Distance,FSD),设计对比学习模块嵌入到包特征到包标签的网络模型,通过端到端的学习优化特征表达和示例分数,提高包标签预测的准确性.在5个数据集上的实验结果表明,本文方法在单数据集实验及数据集间的泛化实验中分别比目前的先进方法平均提高1.16百分点和3.45百分点,尤其在更具挑战性的泛化实验中有明显的优势.
Weakly supervised video anomaly detection methods based on multiple instance learning are currently the dominant approach for anomaly detection.This method leverages instance anomaly scores as the primary learning objective to address the challenge of optimizing the anomaly scores or selection of maximum instance anomaly scores.However,this learning objective tends to downplay the significance of video bag labels.Furthermore,the learning of instance scores is largely constrained by the scarcity of instance labels,thus constraining the effectiveness of anomaly detection.To address the aforementioned challenges,this paper presents an end-to-end video anomaly detection network that integrates fine-grained contrastive learning with video bag label learning as its core learning objective.First,the network utilizes I3D for feature extraction and employs BERT for fea-ture optimization.It then fits the instance scores through a DNN and introduces a Noisy-or model to project the instance scores into bag labels in a way that satisfies multi-instance constraints.This constructs a learning model from bag features to bag labels(Bag-to-bag model),which then enables end-to-end network learning through binary cross-entropy using the true bag labels.Second,to address the inferior instance score learning resulting from the lack of fine-grained annotations,this paper further in-troduces Fine-grained Sequence Distance(FSD)and designs a contrastive learning module embedded into the Bag-to-bag model.Through end-to-end learning,it optimizes the feature representations and instance scores to improve the accuracy of bag label prediction.Experimental results on five datasets show that the proposed approach yields average performance improvements of 1.16 percentage points and 3.45 percentage points over the prevailing state-of-the-art methods in single-dataset experiments and cross-dataset generalization experiments,respectively.Notably,the method shows a distinct advantage in particularly chal-lenging generalization experiments.
丁昕苗;武玉林;郭文;孙昊良
山东工商学院信息与电子工程学院,山东 烟台 264005||中国科学院自动化研究所,北京 100190山东工商学院信息与电子工程学院,山东 烟台 264005山东工商学院信息与电子工程学院,山东 烟台 264005||中国科学院自动化研究所,北京 100190国家互联网应急中心,北京 100029
信息技术与安全科学
异常检测多示例学习Noisy-or模型对比学习深度学习
anomaly detectionmulti-instance learningNoisy-or modelcontrastive learningdeep learning
《计算机与现代化》 2026 (4)
33-40,8
国家自然科学基金资助项目(61876100,62072286)
评论