首页|期刊导航|数字图书馆论坛|基于大语言模型与对比学习的文本可读性自动分级方法研究

基于大语言模型与对比学习的文本可读性自动分级方法研究OA

Research on Automatic Text Readability Grading Method Based on Large Language Model and Contrastive Learning

中文摘要英文摘要

文本可读性是影响用户信息理解、吸收与利用的重要因素.本研究针对现有中文文本可读性自动分级方法在相邻等级区分中容易混淆、高质量标注数据不足等问题,提出了一种基于大语言模型与对比学习的文本可读性自动分级方法.首先利用大语言模型生成具有明确可读性梯度级别的对比样本以扩充训练数据,随后利用BERT模型提取文本的深层语义特征,并引入对比学习策略在特征空间中优化等级边界,以增强模型对相邻等级的判别力.实验结果显示,本研究模型的准确率(0.880)和F1值(0.881)均优于对比基线模型,说明该方法能有效提升中文文本可读性自动分级性能,尤其在相邻等级文本判别方面具有一定优势.

Text readability greatly affects users'information understanding,absorption,and use.To address the problems of confusion between adjacent readability levels and insufficient high-quality labeled data in Chinese text readability grading,this paper proposes an automatic readability grading method based on large language models and contrastive learning.First,a large language model is used to generate contrastive samples with clear readability gradients to expand the training data.Then,BERT is employed to extract deep semantic features,and contrastive learning is introduced to optimize level boundaries in the feature space,thereby improving the model's ability to distinguish adjacent levels.The experimental results demonstrate that the accuracy(0.880)and F1-score(0.881)of the proposed model outperform all baseline models,suggesting that this method can effectively enhance the performance of automatic readability level classification for Chinese texts,particularly showing a clear advantage in discriminating adjacent-level texts.

储伊力;曹振祥;姜烨;兰新苗

安徽医科大学创新创业学院,合肥 230032||公共健康社会治理安徽省哲学社会科学重点实验室,合肥 230032安徽财经大学国际经济贸易学院,蚌埠 233030合肥工业大学计算机与信息学院,合肥 230009中国科学院光电技术研究所,成都 610209

信息技术与安全科学

大语言模型对比学习可读性文本信息自动分级

Large Language ModelContrastive LearningReadabilityText InformationAutomatic Classification

《数字图书馆论坛》 2026 (4)

46-55,10

本研究得到安徽省哲学社会科学规划项目"安徽省老年人健康信息文本可读性评估研究"(编号:AHSKQ2022D141)资助.

10.3772/j.issn.1673-2286.2026.04.005

评论