面向跨语种唇音同步与动态范围增强的真人数字分身生成方法研究OA
Research on digital human generation method oriented to cross-lingual lip-sync and dynamic range enhancement
针对当前真人数字分身在影视级应用中面临的跨语种唇音同步精度低、生成画质动态范围不足等难题,本文提出端到端的全流程解决方案:在语音合成与声纹克隆模块,融合MiniMax-Speech模型与基于检索的语音转换(RVC)变声技术,实现了低资源语言的高保真声纹克隆;在唇音同步模块,通过多语种自适应策略拓展SyncTalk 2D模型对不同语音识别模型的适配范围,提升特殊语种和跨语种情况下的唇形自然度与精准度;在视觉优化模块,引入逆色调映射算法,实现了从标准动态范围(SDR)到符合ITU-R BT.2100标准的高动态范围(HDR)画质转换.实验结果表明,该系统在单张英伟达(NVIDIA)A10显卡环境下推理时长仅为视频总时长的 50%,其图像质量客观评价结果和主观视觉效果优于基线模型.该系统已在新华通讯社新闻播报场景中验证了有效性,可为影视制作、虚拟演播等领域提供技术参考.
To address the challenges faced by photorealistic digital human in cinematic applications,including low precision of cross-lingual lip-syncing and limited dynamic range of generated visuals,this paper proposes a complete end-to-end pipeline.In the speech synthesis and voice cloning module,the MiniMax-Speech model is integrated with retrieval-based voice conversion(RVC)to achieve high-fidelity voice cloning for low-resource languages.In the lip-sync module,a multilin-gual adaptive strategy extends the SyncTalk 2D model's compatibility with various speech recognition models,enhancing naturalness in cross-lingual scenarios.In the visual enhancement module,an inverse tone-mapping algorithm is incorpo-rated to convert standard dynamic range(SDR)video into high dynamic range(HDR)video compliant with the ITU-R BT.2100 standard.Experimental results show that on a single NVIDIA A10 GPU,the inference time is only 50%of the video duration,with objective and subjective quality surpassing baseline.The effectiveness of the system has been vali-dated in news-broadcasting scenarios at Xinhua News Agency,and it can serve as a technical reference for film produc-tion,virtual studio production,and related fields.
百乐夫;张宝亢
新华通讯社通信技术局,北京 100803北京邮电大学数字媒体与设计艺术学院,北京 100876||新华通讯社音视频部,北京 100803
信息技术与安全科学
数字分身多模态算法唇音同步HDR影像生成式人工智能
Digital HumanMultimodal AlgorithmLip SynchronizationHDR CinematicGenerative Artificial Intelligence
《现代电影技术》 2026 (5)
34-41,8
评论