基于大语言模型的用户目标导向小众词汇主题挖掘方法OA
User-targeted rare word topic mining method based on large language models
针对小众主题词挖掘领域黑话语义隐晦、新兴词语料稀缺及语义漂移显著等问题,提出了SeedLLM框架.该方法通过跨上下文一致性约束对黑话术语进行全局推理以刻画其深层语义,并结合检索增强生成补充新兴词外部语义,同时设计新兴词评估模型与在线强化学习机制以抑制过度检索,实现对语义演化的动态适应.多领域数据集实验结果表明,所提方法在小众主题词识别与主题聚合任务上优于现有方法,具有良好的泛化能力与应用价值.
To address challenges in niche topic word mining,including implicit slang semantics,scarce corpora for emerging terms,and significant semantic drift,the SeedLLM framework was proposed.It introduced cross-context con-sistency constraints to perform global reasoning over slang terms and capture their deep semantics.Retrieval-augmented generation was further employed to supplement external semantics for emerging words.Meanwhile,an emerging-word evaluation model and an online reinforcement learning mechanism were designed to suppress excessive retrieval and dy-namically adapt to semantic evolution.The experimental results on multiple cross-domain datasets demonstrate that the proposed method outperforms existing methods in niche topic word identification and topic aggregation,showing strong generalization and practical value.
曾祥;王晔;杨博捷;周斌;王志超;贾焰
国防科技大学计算机学院,湖南 长沙 410073哈尔滨工业大学(深圳)计算机科学与技术学院,广东 深圳 518055国防科技大学计算机学院,湖南 长沙 410073国防科技大学计算机学院,湖南 长沙 410073湖南四方天箭信息科技有限公司,湖南 长沙 410006国防科技大学计算机学院,湖南 长沙 410073
信息技术与安全科学
目标主题挖掘大语言模型强化学习检索增强生成
target topic mininglarge language modelreinforcement learningretrieval-augmented generation
《通信学报》 2026 (7)
16-31,16
国家重点研发计划基金资助项目(No.2023YFC3306103) The National Key Research and Development Program of China(No.2023YFC3306103)
评论