基于纯净韵律表征的表现力语音合成研究OA
Expressive speech synthesis based on pure prosodic representation
针对传统DDPM模型在语音合成中对动态韵律特征捕捉不足、不适配语音特性的问题,提出了一种纯净韵律增强语音合成算法.首先,设计了一种基于分布建模的韵律编码网络,该网络能有效抑制噪声的干扰,学习到纯净的韵律特征嵌入.其次,提出了一种新型扩散模型HarmoVDiff并作为声学模型,比传统DDPM更适配语音特性.最后,进一步构建了一种两阶段语音合成框架HVDIF-TTS,可实现富有表现力和自然度的语音合成.实验结果表明:所提算法在LJSpeech数据集上的BAPD值达 3.681 4,GPE值达 0.150 2,NISQA-MOS值达 3.77.
Aiming at the deficiencies of traditional DDPM models in capturing dynamic prosodic features and their incompatibility with speech characteristics in speech synthesis,this paper proposes an expressive speech synthesis algorithm based on pure prosodic representation.Firstly,a prosody encoding network based on distribution modeling is designed,which can effectively suppress noise interference and learn pure prosodic feature embeddings.Secondly,a new diffusion model,HarmoVDiff,is proposed as the acoustic model,which is more compatible with speech characteristics than traditional DDPMs.Finally,a two-stage speech synthesis framework,HVDIF-TTS,is further constructed to achieve expressive and natural speech synthesis.Experimental results on the LJSpeech dataset show that the proposed algorithm achieves a BAPD value of 3.681 4,a GPE value of 0.150 2,and a NISQA-MOS value of 3.77.
周燕;刘凌志;周月霞;刘翔宇
佛山大学计算机与人工智能学院,广东 佛山 528225佛山大学计算机与人工智能学院,广东 佛山 528225佛山大学计算机与人工智能学院,广东 佛山 528225佛山大学计算机与人工智能学院,广东 佛山 528225
信息技术与安全科学
DDPM模型韵律特征韵律编码网络HarmoVDiff声学模型语音合成
DDPM modelprosodic featuresprosodic encoding networkHarmoVDiff acoustic modelspeech synthesis
《佛山科学技术学院学报(自然科学版)》 2026 (2)
1-9,9
国家自然科学基金资助项目(61972091)
评论