首页|期刊导航|信息通信技术与政策|大语言模型伦理评估中文数据集研究综述

大语言模型伦理评估中文数据集研究综述OA

A review of research on Chinese datasets for ethical evaluation of large language models

中文摘要英文摘要

大语言模型作为人工智能领域的前沿成果,其相关伦理问题备受学界与业界重视,中文语境下的大语言模型伦理评估数据集也随之逐步增多,具备深入研究价值.然而,当前缺乏对这类数据集的系统性梳理与分析,导致研究人员难以精准筛选适配数据集,也无法有效识别现有资源的短板.以2021年8月至2025年3月期间发布的50个中文大语言模型伦理评估数据集为研究对象,从数据集发布时间、创建信息、内容信息、开源情况、涉及领域、伦理场景等方面开展全面对比分析,为后续数据集优化与构建提供方向.

As cutting-edge achievements in the field of artificial intelligence,large language models have drawn significant attention from both academia and industry regarding their associated ethical issues.Consequently,the number of ethical evaluation datasets for large language models in the Chinese context has gradually increased,presenting substantial value for in-depth research.However,the current lack of systematic review and analysis of such datasets makes it difficult for researchers to accurately select suitable datasets and effectively identify shortcomings in existing resources.This paper examines 50 Chinese ethical evaluation datasets for large language models released between August 2021 and March 2025.It conducts a comprehensive comparative analysis covering release dates,creation details,content information,open-source situation,domains covered,and ethical scenarios.This study aims to provide direction for optimizing and constructing future datasets.

田小雨;李文宇;毕春丽;付娜;张蕾蕾

电信科学技术研究院,北京 100191中国信息通信研究院知识产权与创新发展中心,北京 100191中国信息通信研究院知识产权与创新发展中心,北京 100191中国信息通信研究院知识产权与创新发展中心,北京 100191中国信息通信研究院知识产权与创新发展中心,北京 100191

信息技术与安全科学

大语言模型伦理评估中文数据集

large language modelsethical evaluationChinese datasets

《信息通信技术与政策》 2026 (5)

58-68,11

2025年湖南省重大科技攻关项目(No.2025QK2009)

10.12267/j.issn.2096-5931.2026.05.008

评论