首页|期刊导航|湖北民族大学学报(自然科学版)|基于像素级特征调制与文本引导增强的组合零样本学习模型

基于像素级特征调制与文本引导增强的组合零样本学习模型OA

Compositional Zero-shot Learning Model Based on Pixel-level Feature Modulation and Text-guided Refinement

中文摘要英文摘要

针对组合零样本学习(compositional zero-shot learning,CZSL)中未见属性-对象组合泛化能力不足的问题,提出基于像素级特征调制与文本引导增强的组合零样本学习(pixel-level feature modulation and text-guided refinement for compositional zero-shot learning,PFMTR)模型,旨在提升模型对未知组合的识别性能.首先,设计像素级特征调制(pixel-level feature modulation,PLFM)模块,通过像素级与块级特征的双重注意力机制,实现图像特征的精细化重组与语义增强.其次,提出文本引导特征增强(text-guided refinement,TGR)模块,以文本特征为查询、视觉特征为键值,借助跨模态注意力机制计算语义关注权重,实现文本对视觉特征的动态语义引导与跨模态对齐.结果表明,与其他前沿模型相比,PFMTR模型在得克萨斯大学Zappos(University of Texas Zappos,UT-Zappos)数据集上表现突出,曲线下面积(area under curve,AUC)、调和均值(harmonic mean,HM)分别为 35.7%、49.7%.该研究证明通过融合像素级局部特征调制与跨模态语义引导,能够有效增强模型对未见组合的识别性能,为复杂场景下的组合零样本学习提供了可行的技术路径.

To address the insufficient generalization to unseen attribute-object compositions in compositional zero-shot learning(CZSL),a pixel-level feature modulation and text-guided refinement for compositional zero-shot learning(PFMTR)model was proposed to boost the recognition of novel compositions.Firstly,a pixel-level feature modulation(PLFM)module was devised,which employed a dual attention mechanism operating at both pixel-level and patch-level to enable fine-grained reassembly and semantic enhancement of image features.Secondly,a text-guided refinement(TGR)module was proposed,where textual features were used as queries and visual features as keys/values.This module leveraged cross-modal attention to compute semantic relevance weights,thereby dynamically guiding visual features with linguistic semantics and achieving cross-modal alignment.The results showed that,compared with other state-of-the-art models,the PFMTR model achieved outstanding performance on the University of Texas Zappos(UT-Zappos)dataset under the open-world setting,attaining 35.7%in the area under curve(AUC)and 49.7%in the harmonic mean(HM).This study demonstrated that the recognition of unseen compositions could be effectively enhanced by integrating pixel-wise local modulation with cross-modal semantic guidance,offering a viable technical route for CZSL in complex scenarios.

赵薇;包象琳;杜文龙;徐晓峰

安徽工程大学 计算机与信息学院,安徽 芜湖 241000安徽工程大学 计算机与信息学院,安徽 芜湖 241000安徽工程大学 计算机与信息学院,安徽 芜湖 241000安徽工程大学 计算机与信息学院,安徽 芜湖 241000

信息技术与安全科学

组合零样本学习像素级跨模态对齐注意力机制视觉-语言模型

compositional zero-shot learningpixel-levelcross-modal alignmentattention mechanismvision-language model

《湖北民族大学学报(自然科学版)》 2026 (1)

75-81,7

国家自然科学基金项目(62406004)安徽高校自然科学研究项目(2024AH050122)安徽未来技术研究院项目(2023qyhz14).

10.13501/j.cnki.42-1908/n.2026.03.006

评论