面向端到端语音翻译的分阶段训练与策略优化方法OA
A staged training and policy optimization method for end-to-end speech translation
近年来,语音大模型凭借强大的跨模态表征与跨语言理解生成能力,为端到端语音翻译(End-to-End Speech Translation,E2E ST)的研究提供了新范式,现有E2E ST方法也逐渐开始利用大模型的跨模态优势,以缓解传统训练中的对齐与优化难题.然而,直接从语音生成译文仍十分困难,模型需在单一目标下同时完成声学建模、语义理解与跨语言生成,在平行语料有限时,往往容易导致训练不稳定和译文质量下降.为此,提出一种面向端到端语音翻译的分阶段训练与策略优化方法.该方法在统一自回归生成框架下,将关键词、ASR转录和目标译文组织为结构化输出序列,并通过三个阶段逐步提升模型的语音理解与翻译生成能力.首先,通过关键词预测与语音转写联合建模,构建由关键词语义骨架和完整源端转写组成的双粒度源端表示,以建立稳定的声学-语义对应关系;其次,引入完整翻译目标,并结合人工标注与阶段预测结果来构建混合辅助标签,以提升模型对推理阶段中间表示噪声的适应能力;最后,采用基于群体相对策略优化(Group Relative Policy Optimization,GRPO)的强化学习方法,仅针对最终译文子序列进行奖励建模,进一步优化译文生成策略.在Qwen2-Audio,Qwen2.5-Omni和Qwen3-Omni模型上对MuST-C三个语言对进行的实验表明,该方法在BLEU和COMET上均取得显著提升,验证了分阶段训练、混合辅助标签与译文级策略优化能够有效提升端到端语音翻译的稳定性和跨语言生成质量.
In recent years,large speech models have demonstrated strong capabilities in cross-modal representation and cross-lingual generation,offering a new paradigm for end-to-end speech translation(E2E ST).Existing E2E ST methods have gradually leveraged these advantages to alleviate alignment and optimization challenges inherent in traditional training.However,directly generating translations from speech remains difficult,as the model is required to simultaneously perform acoustic modeling,semantic understanding,and cross-lingual generation within a single objective.This challenge becomes more severe when parallel data are limited,often leading to unstable training and degraded translation quality.To address these issues,this paper proposes a staged training and policy optimization method for end-to-end speech translation.Under a unified autoregressive generation framework,the proposed method organizes keywords,ASR transcriptions,and target translations into a structured output sequence and progressively enhances the model's speech understanding and translation generation capabilities through three stages.First,keyword prediction and speech transcription are jointly modeled to construct a dual-granularity source representation comprising a keyword-level semantic skeleton and a complete source transcription,thereby establishing stable acoustic-semantic correspondences.Second,the full translation objective is introduced,and mixed auxiliary labels are constructed by combining human annotations with predictions from the previous stage,improving the model's adaptability to intermediate-representation noise during inference.Finally,reinforcement learning based on Group Relative Policy Optimization(GRPO)is adopted,with reward modeling applied exclusively to the final translation subsequence,to further refine the translation generation strategy.Experiments on three MuST-C language pairs using Qwen2-Audio,Qwen2.5-Omni,and Qwen3-Omni demonstrate that the proposed method achieves significant improvements in both BLEU and COMET scores,confirming that the integration of staged training,mixed auxiliary labels,and translation-level policy optimization effectively enhances both the stability and cross-lingual generation quality of end-to-end speech translation.
朱烨;李军辉;周国栋
苏州大学计算机科学与技术学院,苏州,215006苏州大学计算机科学与技术学院,苏州,215006苏州大学计算机科学与技术学院,苏州,215006
信息技术与安全科学
端到端语音翻译分阶段训练策略优化强化学习
end-to-end speech translationstaged trainingpolicy optimizationreinforcement learning
《南京大学学报(自然科学版)》 2026 (4)
577-591,15
国家自然科学基金(62376178)
评论