基于自然语言处理的职务犯罪法律文书处理与分析研究OA
Research on the processing and analysis of legal documents for duty crimes based on natural language processing
近年来,职务犯罪案件频发,现有研究多局限于法律文本和犯罪构成分析,缺乏跨学科视角,难以揭示其特征和发展趋势.目前,专门针对职务犯罪文书处理与分析的类似系统较少,法律领域通用的数据分析系统难以处理此类文书的专业性和特殊性.因此,借助大数据、人工智能和自然语言处理技术,分析职务犯罪案例文本,揭示犯罪规律并实现高效预防具有重要意义.本研究提出基于智能数据处理与分析的职务犯罪研究模型与算法,并构建了系统原型.通过定制化爬虫技术高效采集多平台职务犯罪文书数据.在数据预处理阶段,采用jieba分词结合深度学习序列标注技术进行清洗、分词及关键信息提取.基于Word2Vec模型将文本信息转化为数字化表达,并结合K-Means聚类算法与Llama3大语言模型挖掘关键特征,显著提升类案检索精准性.最终通过箱线图、散点图等可视化手段展示犯罪规律.实验结果表明,相较于传统方法,该模型在精确度和召回率方面分别提升了21%和9%,充分验证了Llama3在语义理解和特征提取方面的强大能力.
In recent years,there have been frequent cases of job-related crimes,and existing research is mostly limited to legal texts and analysis of crime composition,lacking interdisciplinary perspectives and making it difficult to reveal their characteristics and devel-opment trends.It is of great significance to use big data,artificial intelligence,and natural language processing technologies to analyze case texts of job-related crimes,reveal criminal patterns,and achieve efficient prevention.A research model and algorithm for job-relat-ed crimes based on intelligent data processing and analysis were proposed,and a system prototype was constructed.Efficiently collect multi platform job-related crime document data through customized web crawling technology.In the data preprocessing stage,jieba seg-mentation combined with deep learning sequence annotation technology is used for cleaning,segmentation,and key information extrac-tion.Based on the Word2Vec model,text information is converted into digital expressions,and combined with K-Means clustering algo-rithm and Llama3 big language model to mine key features,significantly improving the accuracy of case retrieval.Finally,crime patterns are displayed through visualization methods such as box plots and scatter plots.The experimental results show that compared to tradition-al methods,the model has improved accuracy and recall by 21%and 9%respectively,fully verifying the powerful ability of Llama3 in se-mantic understanding and feature extraction.
姜志超;杨炳文;高谷刚;李林怡
江苏警官学院 计算机信息与网络安全系,江苏 南京 210031江苏警官学院 计算机信息与网络安全系,江苏 南京 210031江苏警官学院 计算机信息与网络安全系,江苏 南京 210031江苏警官学院 计算机信息与网络安全系,江苏 南京 210031
信息技术与安全科学
职务犯罪法律文书大数据自然语言处理词向量模型聚类算法
Job-related crimesLegal documentsBig dataNatural language processingWord vector modelClustering algorithm
《通信与信息技术》 2026 (1)
7-12,30,7
国家自然科学基金青年项目(项目编号:72401110)
评论