基于EALMDA的医疗命名实体识别数据增强方法OA
EALMDA:a Data Augmentation Method for Medical Named Entity Recognition
医疗命名实体识别是从非结构化医疗文本中识别命名实体,在许多下游任务中起重要作用.医疗命名实体的复杂性需要专家利用领域知识进行标注,导致医疗领域存在严重的标注数据稀缺问题.为解决该问题,提出了一种基于实体感知掩码局部融合命名实体识别数据增强(entity aware mask local mixup data augmentation,EALMDA)方法.首先,使用实体感知掩码通道提取关键元素并掩码非实体部分,以保留核心语义.其次,通过上下文实体相似度和 k近邻两种采样策略的线性组合对掩码句子进行融合,保留核心语义的同时增加样本的多样性.最后,经序列线性化操作后,将句子输入生成的模型中得到增强样本.在 NCBI-disease等五个主流医疗命名实体识别数据集上,模拟低资源场景与主流的数据增强基线方法进行对比实验,所提方法的性能相比基线方法有显著提升.
Named entity recognition in the medical field involved identifying named entities from unstruc-tured medical texts.This played a crucial role in various downstream tasks.Due to the complexity of medical named entities leveraging domain-specific knowledge was required in expert annotations,which led to a severe scarcity of annotated data in the medical field.To address this issue,an entity-aware mask local mixup data augmentation method(EALMDA)was proposed.This method firstly extracted key ele-ments using an entity-aware masking channel and masked non-entity parts to retain core semantics.Then,masked sentences were fused through a linear combination of two sampling strategies:contextual entity similarity and k-nearest neighbors,which preserved the core semantics while increasing sample di-versity.Finally,after sequence linearization,the sentences were input into a generative model to obtain augmented samples.Comparative experiments were conducted on five mainstream medical named entity recognition datasets,such as NCBI-disease,simulating low-resource scenarios against mainstream data augmentation baselines,and significant improvements were observed compared to baseline methods.
道路;刘纳;郑国风;李晨;杨杰
北方民族大学 计算机科学与工程学院 宁夏 银川 750021||北方民族大学 图像图形智能处理国家民委重点实验室 宁夏 银川 750021北方民族大学 计算机科学与工程学院 宁夏 银川 750021||北方民族大学 图像图形智能处理国家民委重点实验室 宁夏 银川 750021北方民族大学 计算机科学与工程学院 宁夏 银川 750021||北方民族大学 图像图形智能处理国家民委重点实验室 宁夏 银川 750021北方民族大学 计算机科学与工程学院 宁夏 银川 750021||北方民族大学 图像图形智能处理国家民委重点实验室 宁夏 银川 750021北方民族大学 计算机科学与工程学院 宁夏 银川 750021||北方民族大学 图像图形智能处理国家民委重点实验室 宁夏 银川 750021
信息技术与安全科学
数据增强命名实体识别自然语言处理生成模型Mixup
data augmentationnamed entity recognitionnatural language processinggenerative mod-elMixup
《郑州大学学报(理学版)》 2026 (1)
43-50,8
国家自然科学基金项目(62162001)宁夏自然科学基金项目(2021AAC03224)北方民族大学2024年度校级一般科研项目(2024XYZJK01)
评论