XBMU-AMDO31:藏语安多方言语音识别数据集OA
XBMU-AMDO31:A dataset of speech recognition dataset for the Amdo dialect of Tibetan
近年来,尽管语音识别技术在高资源语种(如英语、汉语)中取得显著进展,但针对藏语等低资源复杂音系语种的研究进展仍然缓慢.安多藏语作为低资源复杂音系语言,其语音识别面临数据稀缺与可用数据集质量和多样性不足的双重挑战.由于缺乏公开的数据集,相关研究面临着诸多限制.为此,本文构建并介绍了一个开源的藏语安多方言语音识别数据集.语音样本最初采集于中国甘肃省夏河地区,共收录了66位以藏语为母语者共31小时录音以及相应的转录文本,后续经过人工质检与标准化处理,确保了方言纯正性的以及数据的质量和一致性.本语音数据集的所有资源均已开放,目前已在多篇藏语语音识别相关论文或研究中被使用,得到业内专家的一致好评,更证明了数据集的质量.本数据集为藏语安多方言的高质量语音数据提供了重要补充,其复杂音系特性为跨语种迁移学习、小样本语音技术研究提供独特样本支持.
In recent years,while significant progress has been made in speech recognition technology for high-resource languages(such as English and Mandarin),research on low-resource languages with complex phonologies,like Tibetan,has progressed relatively slow.As a low-resource and phonologically complex language,Amdo Tibetan faces dual challenges in speech recognition:data scarcity and insufficient quality and diversity of available datasets.The lack of publicly accessible datasets has imposed numerous constraints on related research.To address these challenges,this paper introduces and presents an open-source speech recognition dataset for the Amdo Tibetan dialect.The speech samples were initially collected in Xiahe County,Gansu Province,China,comprising 31 hours of recordings from 66 native speakers along with corresponding transcriptions.Subsequent manual quality control and standardization were applied to ensure the authenticity of the dialect as well as the consistency and quality of the data.All resources in this dataset have been made publicly available and have already been utilized in multiple research papers and studies on Tibetan speech recognition,receiving widespread acclaim from experts in the field—further validating the dataset's quality.This dataset serves as an important supplement to high-quality speech data for Amdo Tibetan,and provides unique support for cross-lingual transfer learning and few-shot speech technology research due to its complex phonological characteristics.
谢晨宇;李冠宇;马立克;孙倩;郭玉豪
西北民族大学语言与文化计算教育部重点实验室,兰州 730030西北民族大学语言与文化计算教育部重点实验室,兰州 730030||青藏高原人文环境数据智能实验室,兰州 730030西北民族大学语言与文化计算教育部重点实验室,兰州 730030西北民族大学语言与文化计算教育部重点实验室,兰州 730030西北民族大学语言与文化计算教育部重点实验室,兰州 730030
语音识别安多藏语数据集多说话人低资源
speech recognitionAmdo Tibetan datasetmulti-speakerlow-resource
《中国科学数据(中英文网络版)》 2026 (1)
43-53,11
国家自然科学基金(61633013)2024年甘肃省科技重大专项计划项目(24ZDFA004). National Natural Science Foundation of China(61633013)Gansu Provincial Science and Technology Major Project(24ZDFA004)
评论