基于多任务预训练的情感音乐生成方法OA
Multi-task Pre-training Framework for Emotional Music Generation
基于人工智能的情感音乐生成(AI-AMG)具备根据环境因素、听众生理或情感状态的变化灵活定制生成音乐的能力,在心理健康疗愈、沉浸式声像表现、智能音乐创作等多个领域展现出巨大影响力和潜力.由于情感音乐标注任务成本高,情感音乐数据集稀缺,从而现有情感音乐生成方法缺乏足够的有效数据进行学习,模型的情感感知能力较弱,生成音乐的情感表达准确性不足.提出一种面向情感音乐理解和生成的多任务预训练-微调框架,旨在提升生成音乐的情感表达准确性.预训练阶段基于大量无标签数据设计与情感特征关联的小节属性掩码任务、调式预测任务与旋律补全任务,使用不同注意力机制实现音乐特征与情感语义的表示学习;微调阶段结合带情感标注的音乐数据集,引入效价-唤醒度情感属性的八元组表示方法,应用序列到序列的注意力机制实现情感可控的音乐生成.实验表明,所提方法相较于基准模型四分类准确率提升4.2个百分点,二分类准确率则分别提升17.48和3.40个百分点,在主观评估中,表现出更高的听感质量,消融实验进一步验证了各预训练子任务对情感建模的有效性.
AI-based affective music generation(AI-AMG)enables flexible music creation tailored to environmental contexts and listeners' physiological or emotional states,aiming to stimulate and regulate affective responses.It demonstrates strong potential in domains such as mental health therapy,immersive multimedia experiences,and intelligent music composition.Challenges remain,including the scarcity of affective music datasets,the difficulty of emotion annotation,the limited perceptual capability of existing models,and low accuracy in affective expression due to ineffective learning paradigms.A unified framework for affective music understanding and generation is proposed by integrating multi-task pre-training with fine-tuning.During the pre-training phase,proxy tasks,including bar-level attribute masking,mode prediction,and melody completion are designed to associate emotional characteristics with structural musical features using large-scale unlabeled data.Task-specific attention mechanisms enable joint learning of musical and affective representations.In the fine-tuning phase,an octuple representation incorporating valence-arousal dimensions is introduced,allowing emotion-controllable music generation on annotated datasets via a sequence-to-sequence attention mechanism.Experimental results show that the method improves four-class emotion classification accuracy by 4.2 percentage points and achieves gains of 17.48 and 3.40 percentage points in two-class classification tasks compared with baseline model.Subjective evaluation confirms enhanced perceived audio quality,and ablation experiments further validate the effectiveness of each pre-training sub-task for sentiment modeling.
范舒柯;王海龙;柳林
内蒙古师范大学 计算机科学技术学院,呼和浩特 010022内蒙古师范大学 计算机科学技术学院,呼和浩特 010022||内蒙古师范大学 科学技术史研究院,呼和浩特 010022内蒙古师范大学 计算机科学技术学院,呼和浩特 010022
信息技术与安全科学
音乐信息检索(MIR)情感音乐生成多任务预训练音乐表示
music information retrieval(MIR)emotional music generationmulti-task pre-trainingmusic representation
《计算机科学与探索》 2026 (7)
2027-2036,10
国家自然科学基金(62566047)国家重点研发计划(2020YFC1523305)内蒙古自治区自然科学基金(2024LHMS06015)2022年度国家社科基金冷门绝学研究专项学术团队项目(22VJXT008). This work was supported by the National Natural Science Foundation of China(62566047),the National Key Research and Development Program of China(2020YFC1523305),the Natural Science Foundation of Inner Mongolia Autonomous Region(2024LHMS06015),and the 2022 Academic Team Project of the National Social Science Fund of China Special Program for Endangered and Obscure Disciplines(22VJXT008).
评论