大模型个人信息训练语料的合法性困境与标准再造OA
The legitimacy dilemma and standard reconstruction of personal information training corpora for large language models
个人信息是大模型训练语料的重要组成部分,事关其推理能力的养成.现有个人信息保护规范遵循基于权利的方法,因无法适应大模型黑箱、数据挖掘、披露制度等技术特征,存在主体地位不平等、权利价值冲突与防御被动等局限.为此,有必要转向基于风险的方法,通过知情同意的弱化适用、过程控制的主体转移、彻底删除的程度消减三方面的标准设定,克服既往规范的滞后性,以实现风险责任合理分配、治理效率提升、多元价值平衡,最终呈现大模型训练的全流程个人信息保护方案.
Personal information constitutes a vital component of training corpora for large language models(LLMs),as it is crucial to the de-velopment of their reasoning capabilities.Existing personal information protection norms are based on a rights-oriented approach,which strug-gles to adapt to the technical characteristics of LLMs—such as algorithmic black boxes,data mining,and disclosure mechanisms—and thus suffers from limitations including unequal subject status,conflicts in rights values,and passive defense measures.To address these issues,it is imperative to shift to a risk-oriented approach.Guided by this approach,establishing standards in three key areas—the flexible adaptation of informed consent,the transfer of process control responsibilities,and the mitigation of complete deletion requirements—can overcome the lag of previous norms.This shift enables the rational allocation of risk liabilities,enhances governance efficiency,and balances diverse values,ulti-mately forming a full-lifecycle personal information protection framework for LLM training.
葛佳怡
山东大学 法学院,山东 青岛 266237
社会科学
大模型个人信息保护信息安全基于权利的方法基于风险的方法
large language modelpersonal information protectioninformation securityrights-based approachrisk-based approach
《网络安全与数据治理》 2026 (8)
77-83,7
评论