首页|期刊导航|大数据|大语言模型的自我恒定性现象及其在知识增强任务上的应用

大语言模型的自我恒定性现象及其在知识增强任务上的应用OA

Self-consistency phenomenon in large language model and its application in knowledge augmentation tasks

中文摘要英文摘要

近期研究指出,检索和大语言模型中存在显著的"源偏见",即相较于真实人类数据,其更偏好模型生成的内容.人类存在类似的"自我恒定性"现象,即倾向于维护自我概念的一致性,相较于外部事实,其更倾向于相信已有的知识.探讨大模型是否表现出与类人的认知偏差,分别从显式方式(模型自身生成的内容)与隐式方式(提示词诱导模型认为内容是其生成的)两方面展开研究,并结合自评估与一致性两种置信度评价方法,在GPT-4o、DeepSeek-R1、Llama2(7B~70B)、Qwen2(7B~72B)等模型上进行了系统实验.结果表明,在隐式方式下,模型表现出明显的自我恒定性,对自身内容有显著更高的信心.利用该特性可以系统提升模型在各类知识增强任务中的表现:在TriviaQA、NQ、HotpotQA、FEVER和ZsRE等数据集上,角色提示策略带来准确率提升;在存在大量噪声干扰文档时,仍能保持稳健优势,显示出良好的鲁棒性和泛化性.最后,揭示了指令微调与人类反馈强化学习是导致模型产生此类偏见的核心因素,从训练与对齐机制层面解释了自我恒定性的形成.

Recent studies have revealed significant"source bias"in retrieval and large language models(LLM),where these models exhibit a preference for model-generated content over authentic human-written data.Humans similarly demonstrate"self-consistency"phenomena,characterized by a tendency to maintain coherent self-concepts and favor existing knowledge over external facts.It was investigated whether LLMs exhibit human-like cognitive biases.It was conducted from two perspectives:an explicit approach(content directly generated by the LLM)and an implicit approach(using prompts to induce the models to perceive the content as self-generated).Confidence was evaluated using both self-assessment and consistency methods.Systematic experiments were carried out on several models,including GPT-4o,DeepSeek-R1,Llama2(7B-70B),and Qwen2(7B-72B).The results demonstrated that under implicit conditions,LLMs exhibit pronounced self-consistency phenomena,displaying significantly higher confidence in self-attributed content.Leveraging this characteristic yields consistent performance gains:the proposed role-based prompting strategy improved accuracy on TriviaQA,NQ,HotpotQA,FEVER,and ZsRE,and maintained a clear advantage under high levels of noisy retrieved documents,indicating strong robustness and generalization.Finally,our mechanistic analysis revealed that instruction fine-tuning and reinforcement learning from human feedback were the primary factors contributing to such biases in models,explaining how self-consistency emerges during training.

雷艳;庞亮;魏子豪;王元卓;沈华伟;程学旗

中国科学院计算技术研究所智能算法安全全国重点实验室,北京 100190||中国科学院大学计算机学院,北京 100049中国科学院计算技术研究所智能算法安全全国重点实验室,北京 100190中国科学院计算技术研究所智能算法安全全国重点实验室,北京 100190||中国科学院大学计算机学院,北京 100049中国科学院计算技术研究所智能算法安全全国重点实验室,北京 100190||中国科学院大学计算机学院,北京 100049中国科学院计算技术研究所智能算法安全全国重点实验室,北京 100190||中国科学院大学计算机学院,北京 100049中国科学院计算技术研究所智能算法安全全国重点实验室,北京 100190||中国科学院大学计算机学院,北京 100049

信息技术与安全科学

大语言模型检索增强知识增强任务源偏见

large language modelretrieval-augmented generationknowledge augmentationsource bias

《大数据》 2026 (3)

98-114,17

国家自然科学基金资助项目(No.62172393)河南省重大公益项目(No.201300311200) The National Natural Science Foundation of China(No.62172393),The Major Public Welfare Project of Henan Province(No.201300311200)

10.11959/j.issn.2096-0271.2026038

评论