XBMU-bo-Lhasa31:藏语拉萨话语音识别数据集OA
XBMU-bo-Lhasa31:A dataset of speech recognition for the Lhasa Dialect of Tibetan
藏语语音识别在藏语教育、新闻传播等领域具有重要应用价值.藏语拉萨话广泛使用于拉萨市及周边地区,由于地域等因素的影响,当前可用的藏语语音数据资源匮乏,高质量标注数据稀缺.为此,本研究构建了一个专业规范的藏语拉萨话语音识别数据集.数据集使用自制录音软件实地录制,采集自51位说话人,总时长31.61小时,包含24,289条语音样本,平均每条时长4.68秒.数据内容主要选自新闻领域文本,确保语言规范性和领域代表性.为保障数据质量,实施了严格的质量控制流程:首先,对原始文本进行分句处理和人工校验;其次,在录音完成后,采用语音端点检测(VAD)技术筛选优质录音样本;最后,对文本中的非发音符号进行规范化处理,以提高语音识别的准确性.本数据集的建立为藏语语音识别研究提供了重要基础资源,对推动藏语语音识别技术发展具有积极意义.
Tibetan speech recognition has important application value in fields such as Tibetan language education,news dissemination and other fields.The Lhasa dialect of Tibetan is widely used in Lhasa City and its surrounding regions.However,due to geographical and other constrains,currently available Tibetan speech data resources remained limited and high-quality annotated data are particularly scarce.For this reason,this study constructs a professionally designed and standardized speech recognition dataset for the Lhasa dialect of Tibetan.The dataset was recorded in real-world environments using self-developed recording software,and was collected from 51 speakers,with a total duration of 31.61 hours,containing 24,289 speech samples,with an average duration of 4.68 seconds per sample.The data content was primarily selected from news-related texts to ensure linguistic standardization and domain representativeness.In order to guarantee data quality,we implemented a strict quality control process:firstly,the original texts were segmented into sentences and manually verified;after the recordings were completed,the Voice Activity Detection(VAD)technique was used to filter and regain high-quality speech samples;in addition,non-pronounced symbols in the text were normalized to improve the accuracy of speech recognition.The establishment of this dataset provides an important foundational resource for Tibetan speech recognition and is expected to facilitate the development of Tibetan speech recognition technology.
马立克;李冠宇;谢晨宇;孙倩;郭玉豪
西北民族大学语言与文化计算教育部重点实验室,兰州 730030西北民族大学语言与文化计算教育部重点实验室,兰州 730030||青藏高原人文环境数据智能实验室,兰州 730000西北民族大学语言与文化计算教育部重点实验室,兰州 730030西北民族大学语言与文化计算教育部重点实验室,兰州 730030西北民族大学语言与文化计算教育部重点实验室,兰州 730030
语音识别藏语拉萨话多说话人语音语料库
Speech RecognitionTibetan Lhasa dialectMulti-Speakerphonetic corpus
《中国科学数据(中英文网络版)》 2026 (1)
31-42,12
国家自然科学基金(61633013)2024年甘肃省科技重大专项计划(24ZDFA004). National Natural Science Foundation of China(61633013)Gansu Province Science and Technology Major Special Program in 2024(24ZDFA004).
评论