融合音频文本的多模态音乐情感联合学习识别OA
Multi-modal music emotion joint learning recognition method integrating audio text
针对音乐多模态情感识别问题,提出一种考虑多模态数据特征的联合学习框架,将多模态情感识别作为主任务,多模态情感类别识别作为辅助任务,通过情感类别识别任务辅助情感识别任务,实现音乐情感与情感类别识别性能共同提升.首先,通过私有层网络对主任务中的音频、文本信息分别进行编码,对单音频、文本模态信息的内部情感特征进行学习.然后,通过共享网络层对主任务中的情感信息以及辅助任务中的情感分类信息进行学习,再分别与主任务中的单模态独立特征相结合,得到主任务上单模态信息.最后,通过自注意力机制捕捉多模态特征信息,获取最终的情感与情感类别识别.实验表明,论文所提方法能够通过联合学习框架提升情感识别性能,同时情感类别识别性能也有一定提升.
Aiming at the problem of multi-modal emotion recognition in music,a joint learning framework considering the characteristics of multi-modal data was proposed.The multi-modal emotion recognition was taken as the main task,and the multi-modal emotion category recognition was taken as the auxiliary task.Through the emotion category recognition task assisting emotion recognition task,the performance of music emotion and emotion category recognition was improved together.Firstly,the audio and text information in the main task were encoded respectively through the private layer network,and the internal emotional characteristics of single audio and text modal information were learned.Then,through the shared network layer,the emotion information in the main task and the emotion classification information in the auxiliary task were learned,and then combined with the single mode independent features in the main task.The single mode information on the main task was obtained.Finally,multi-modal feature information was captured by self-attention mechanism to obtain the final emotion and emotion category recognition.Experiments showed that the proposed method can improve the performance of emotion recognition through the framework of joint learning,and the performance of emotion category recognition can also be improved to some extent.
刘羚瑶
山西传媒学院表演学院,山西晋中 030619
信息技术与安全科学
音乐多模态情感识别联合学习
musicmulti-modalemotion recognitionjoint learning
《安徽大学学报(自然科学版)》 2026 (2)
25-33,9
山西省自然科学基金面上项目(202403021221193)教育部人文社会科学研究规划基金资助项目(23YJAZH070)山西省艺术科学规划课题(22BA152)
评论