层级特征融合Transformer的图像分类算法OA
Image Classification Algorithm Based on Hierarchical Feature Fusion Transformer
针对传统ViT(Vision Transformer)模型难以完成图像多层级分类问题,文中提出了基于ViT的图像分类模型层级特征融合视觉Transformer(Hierarchical Feature Fusion Vision Transformer,HICViT).输入数据经过ViT提取模块生成多个不同层级的特征图,每个特征图包含不同层次的抽象特征表示.基于层级标签将ViT提取的特征映射为多级特征,运用层级特征融合策略整合不同层级信息,有效增强模型的分类性能.在CIFRA-10、CIFRA-100 和CUB-200-2011 这 3 个数据集将所提模型与多种先进深度学习模型进行对比和分析.在CIFRA-10 数据集,所提方法在第 1 层级、第2 层级和第 3 层级的分类精度分别为 99.70%、98.80%和 97.80%.在CIFRA-100 数据集,所提方法在第 1 层级、第 2层级和第 3 层级的分类精度分别为 95.23%、93.54%和 90.12%.在CUB-200-2011 数据集,所提方法在第 1 层级和第 2层级的分类精度分别为 98.09%和 93.66%.结果表明,所提模型的分类准确率优于其他对比模型.
In view of the problem that the traditional ViT(Vision Transformer)model is difficult to complete multi-level image classification,this study proposes a HICViT(Hierarchical Feature Fusion Vision Transformer)for image classification based on ViT.The input data is processed through the ViT extraction module to generate multiple feature maps at different levels,and each feature map contains abstract feature representations at different levels.Ac-cording to the hierarchical labels,the features extracted by ViT are mapped into features at different levels,and a HIC method is used to fuse the features at different levels,thereby improving the classification performance of the model.The proposed model is compared and analyzed with a variety of advanced deep learning models on three data-sets,namely CIFRA-10,CIFRA-100,and CUB-200-2011.On the CIFRA-10 dataset,the classification accura-cies of the proposed method at the first level,the second level,and the third level are 99.70%,98.80%,and 97.80%,respectively.On the CIFRA-100 dataset,the classification accuracies of the proposed method at the first level,the second level,and the third level are 95.23%,93.54%,and 90.12%,respectively.On the CUB-200-2011 dataset,the classification accuracies of the proposed method at the first level and the second level are 98.09%and 93.66%,respectively.The results indicate that the classification accuracy of the proposed model outperforms that of other comparative models.
DUAN Shixi;WANG Bo
School of Science,Shenyang University of Technology,Shenyang 110870,ChinaSchool of Science,Shenyang University of Technology,Shenyang 110870,China
信息技术与安全科学
深度学习卷积神经网络Transformer图像分类层级特征特征融合多头注意力Vision Transformer
deep learningconvolutional neural networksTransformerimage classificationhierarchical charac-teristicsfeature fusionmulti-head attentionVision Transformer
《电子科技》 2026 (2)
72-78,7
国家自然科学基金(62103289)National Natural Science Foundation of China(62103289)
评论