首页|期刊导航|现代信息科技|基于BERT微调与领域词典的简历信息抽取

基于BERT微调与领域词典的简历信息抽取OA

Resume Information Extraction Based on BERT Fine-tuning and Domain Dictionary

中文摘要英文摘要

为了在小样本训练集情况下提升简历中结构化信息抽取的综合效果,为简历信息高效筛选提供可行途径,文章在领域标注数据稀缺的情况下,通过构建竞赛专业领域词典并利用其对少量简历进行数据增广,生成高质量训练语料.基于 BERT 模型,在不同规模增广数据上训练后进行关键信息抽取并计算评价指标.实验结果表明,针对两个测试集进行信息抽取的 F1 值分别达到 91.79%和 76.47%,验证了"领域词典+数据增广"策略在低资源命名实体识别任务中的有效性.

In order to enhance the comprehensive performance of structured information extraction from resumes under small training sample conditions,and provide a viable approach for efficient resume screening,this paper proposes a method to generate high-quality training corpora in the absence of domain-specific annotation data by constructing a competition-specific domain dictionary and leveraging it to augment datasets from a limited number of resumes.Based on the BERT model,key information extraction is performed after training on datasets of varying augmentation scales,and evaluation metrics are calculated.The experimental results indicate that the F1 scores for information extraction on two test sets reach 91.79%and 76.47%,validating the effectiveness of the"domain dictionary+data augmentation"strategy in low-resource named entity recognition tasks.

钱华远

河北工业大学 人工智能与数据科学学院,天津 300401

信息技术与安全科学

命名实体识别自然语言处理领域词典数据增广BERT模型信息抽取

Named Entity RecognitionNatural Language Processingdomain dictionarydata augmentationBERT modelinformation extraction

《现代信息科技》 2026 (11)

68-72,80,6

10.19850/j.cnki.2096-4706.2026.11.013

评论