XBMU-MC:多语言平行语料库OA
XBMU-MC:A Multilingual Parallel Corpus
多语言平行语料库是自然语言处理领域突破语言壁垒、提升跨语言任务效率和质量的基础资源.针对当前汉藏、汉维、汉蒙等低资源语言平行语料相对稀缺的现状,本研究构建了一个面向多语言机器翻译和跨语言信息检索任务的平行语料库——Northwest Minzu University-Multilingual Corpus(XBMU-MC).原始语料包括人工构造特定双语文本和网络爬取公开多语言数据.人工构造语料由研究人员围绕文化、科技、社会等领域撰写汉语文本,并通过机器翻译与人工校对获得目标语言译文;网络爬取语料则从主流藏文、蒙文和维文新闻网站抓取原始文本,扩大语料规模和领域覆盖范围.之后采用数据增强技术对现有语料进行扩展,并对数据进行严格筛选,确保每对平行语料在源语言和目标语言之间的对齐一致性与翻译准确性.最终筛选出21,579条高质量样本,以JSON格式存储,包括指令、输入和输出三个属性.最后,基于GLM4-9B、Qwen3-8B和Deepseek-R1-8B模型对XBMU-MC数据集与开源MMDS数据集进行多语言翻译对比实验,在BLEU得分上分别平均提升21.16、20.51和15.37个百分点.本数据集能在一定程度上缓解低资源语言平行语料不足的问题,可有效支持模型训练与多语言翻译任务,也可作为多语种大语言模型在跨语言对齐、指令微调及民族语言智能信息处理应用的语料基础,助力相关任务性能提升.
Multilingual parallel corpus is a basic resource for natural language processing to break through language barriers and improve the efficiency and quality of cross-language tasks.To address the relative scarcity of parallel corpora for low-resource languages(Chinese-Tibetan,Chinese-Uyghur,Chinese-Mongolian),this study constructs a parallel corpus—XBMU-MC(Northwest Minzu University-Multilingual Corpus)for multilingual machine translation and cross-language information retrieval tasks.The original corpus consists of manually constructed specific bilingual texts and web-crawled publicly available multilingual data.The manually constructed corpus consists of Chinese texts written by researchers in the fields of culture,science and technology,and society,and the target language translations are obtained through machine translation and manual proofreading;while the web-crawled corpus captures the original texts from mainstream Tibetan,Mongolian,and Uyghur news websites to expand the scale of the corpus and the scope of the domains covered.After that,data enhancement techniques are used to expand the existing corpus,and the data are strictly screened to ensure the alignment consistency and translation accuracy between the source and target languages for each parallel corpus pair.Finally,21,579 high-quality samples were selected and stored in JSON format,including three attributes:instruction,input and output.Finally,the multilingual translation comparison experiments between the XBMU-MC dataset and the open-source MMDS dataset based on the GLM4-9B,Qwen3-8B,and Deepseek-R1-8B models show an average increase in BLEU scores of 21.16,20.51,and 15.37 percentage points.This dataset can help mitigate the problem of insufficient parallel corpus for low-resource languages to a certain extent,effectively support model training and multilingual translation tasks,and can also be used as a corpus basis for cross-lingual large language models in cross-language alignment,instruction fine-tuning and intelligent information processing applications in ethnic languages,helping to improve the performance of related tasks.
严琦栋;马宁;巴桑珠扎;白玛曲扎;艾科拜·依米提;麦迪那木·吾斯曼;木巴热克·阿布力克木;苏日娜;阿力娅
西北民族大学,语言与文化计算教育部重点实验室,兰州 730030||西北民族大学,甘肃省民族语言文化智能信息处理重点实验室,兰州 730030西北民族大学,语言与文化计算教育部重点实验室,兰州 730030||西北民族大学,甘肃省民族语言文化智能信息处理重点实验室,兰州 730030西北民族大学,数学与计算机科学学院,兰州 730030西北民族大学,数学与计算机科学学院,兰州 730030西北民族大学,数学与计算机科学学院,兰州 730030西北民族大学,数学与计算机科学学院,兰州 730030西北民族大学,数学与计算机科学学院,兰州 730030西北民族大学,中国语言文学学部,兰州 730030西北民族大学,中国语言文学学部,兰州 730030
汉藏汉蒙汉维平行语料库数据增强
Chinese-TibetanChinese-MongolianChinese-Uyghurparallel corpusdata enhancement
《中国科学数据(中英文网络版)》 2026 (1)
68-81,14
国家自然科学基金项目(62466052)甘肃省中央引导地方科技发展资金项目(25ZYJA034)西北民族大学中央高校基本科研业务费专项资金项目(31920250008) National Natural Science Foundation of China Project(62466052)Gansu Province Central Guidance Local Science and Technology Development Fund Project(25ZYJA034)Fundamental Research Funds for the Central Universities of Northwest Minzu University(31920250008).
评论