知识增强与协同推理的视觉常识推理方法OA
Knowledge-augmented collaborative reasoning for visual commonsense reasoning
视觉常识推理(Visual Commonsense Reasoning,VCR)旨在让模型像人类一样基于图像进行认知推理,现有方法主要依赖大规模跨模态预训练模型进行隐式跨模态建模,在外部知识引入、上下文建模以及答案与理由一致性方面仍存在不足.针对上述问题,本文提出一种知识增强与协同推理的视觉常识推理框架.该框架通过引入由场景常识图和区域事实描述构成的双层知识注入机制,显式融合场景级常识与实体级视觉事实,从而缓解知识缺失与语义模糊问题;同时,构建基于大语言模型的双任务协同学习框架,联合优化问答与理由推断任务,并通过认知对齐损失增强二者在隐空间中的逻辑一致性.在 VCR 数据集上的实验结果表明,所提出的框架能够有效提升模型的推理性能,其中在联合任务 Q→AR 上相比当前代表性方法提升 4.2 个百分点,验证了知识增强与协同推理策略的有效性.
Visual Commonsense Reasoning(VCR)aims to enable models to perform human-like cognitive reasoning based on images.Although existing methods primarily rely on large-scale cross-modal pre-trained models for implicit cross-modal modeling,they still fall short in terms of external knowledge incorporation,contextual modeling,and the consistency between answers and rationales.To address these limitations,the paper propose a knowledge-augmented collaborative reasoning framework for VCR.The proposed frame-work introduces a dual-layer knowledge injection mechanism comprising scene commonsense graphs and re-gional factual descriptions.By explicitly integrating scene-level commonsense with entity-level visual facts,this mechanism mitigates knowledge deficiency and semantic ambiguity.Meanwhile,a dual-task collabora-tive learning framework based on large language models is designed to jointly optimize question answering and rationale inference.A cognitive alignment loss is further employed to enhance the logical consistency be-tween the two tasks in the latent space.Experimental results on the VCR benchmark demonstrate that the proposed framework effectively improves model reasoning performance.Notably,it outperforms state-of-the-art baselines by 4.2 percentage points on the joint Q→AR task,thereby validating the effectiveness of the knowledge-augmented collaborative reasoning strategy.
李星悦;韩春燕;鲜雨成;琚生根;李勤
四川大学计算机学院,成都 610065四川民族学院智能科学与技术学院,康定 626001四川大学计算机学院,成都 610065四川大学计算机学院,成都 610065四川大学计算机学院,成都 610065
信息技术与安全科学
大语言模型视觉常识推理知识增强协同学习
Large language modelsvisual commonsense reasoningknowledge enhancementcollaborative learning
《四川大学学报(自然科学版)》 2026 (4)
862-876,15
国家自然科学基金重点项目(62137001)
评论