首页|期刊导航|山西大学学报(自然科学版)|基于图文理解增强的教科书视觉问答方法

基于图文理解增强的教科书视觉问答方法OA

Enhancing Image and Text Comprehension for Textbook Visual Question Answering

中文摘要英文摘要

教科书视觉问答是智慧教育领域的一项多模态任务,需要深度理解教科书图像、文本和问题以推理正确答案.然而,现有通用领域视觉问答方法在这一任务中表现不佳.主要原因是:首先,这些方法仅能简单地识别物体属性,缺乏学科信息,且容易受到与问题无关的冗余信息干扰;其次,这些方法难以捕捉文本关键信息.针对上述问题,提出基于图文理解增强的教科书视觉问答方法,主要包括3个模块:(1)文本编码与理解:利用大语言模型提取问题关键词,并在文本中检索与问题关键词相关语句,来增强文本理解,消除冗余信息干扰.(2)图像编码与描述:根据问题关键词,在图像描述中采用问题-图像注意力机制生成问题约束的细粒度图像描述语句,从而增强图像理解能力.(3)答案预测:使用预训练视觉语言模型,将文本信息与视觉信息融合,以提高模型推理能力.实验结果表明:该文提出的方法能够较好地理解教科书图文信息,有效提升答案预测准确率.在测试集和验证集上准确率分别提高了1.82%和1.72%.

Textbook Visual Question Answering is a multi-modal task in the field of smart education that requires a deep understand-ing of textbook images,text,and questions to infer the correct answers.However,existing generic Visual Question Answering meth-ods perform poorly in this task.The main reasons are as follows:Firstly,these methods can only simply recognize object attributes,lack disciplinary information,and are susceptible to interference from redundant information unrelated to the questions.Secondly,they struggle to capture key information in the texts.To solve these problems,a textbook visual question-answering method based on image description enhancement is proposed,which mainly includes three modules:(1)Text encoding and understanding:Utilizing large language models to extract keywords from questions and retrieve relevant statements in the text related to the question key-words to enhance text understanding and eliminate interference from redundant informations.(2)Image encoding and description:Employing a question-image attention mechanism in image descriptions to generate fine-grained image description statements con-strained by questions based on question keywords,thereby enhancing image understanding ability.(3)Answer prediction:using a pre-trained visual-language model to fuse text information with visual information to improve the model's reasoning ability.Experi-mental results on relevant datasets demonstrate that the proposed method effectively improves the understanding of textbook infor-mation,thereby enhancing answer prediction accuracy.The accuracy of the test set and the verification set was improved by 1.82%and 1.72%,respectively.

胡景畅;强鹏鹏;谭红叶;王宏宇;慕永利

山西大学 计算机与信息技术学院,山西 太原 030006山西大学 计算机与信息技术学院,山西 太原 030006山西大学 计算机与信息技术学院,山西 太原 030006||计算智能与中文信息处理教育部重点实验室,山西 太原 030006山西大学 计算机与信息技术学院,山西 太原 030006智林信息技术股份有限公司,山西 太原 030000

信息技术与安全科学

视觉问答智慧教育图像描述图文理解增强

visual question answeringintelligent educationimage captionimage-text comprehension enhancement

《山西大学学报(自然科学版)》 2026 (2)

263-271,9

国家自然科学基金(62076155)太原市小店区-山西大学产学研合作项目"短答案自动评分技术在综合评价系统中的推广与应用"(202301S06)

10.13451/j.sxu.ns.2024104

评论