基于双标签混合Transformer的骨骼关键点重定向方法OA
Skeleton keypoints retargeting method based on dual-label hybrid Transformer
针对人体动作迁移中骨骼比例差异导致的图像生成失真问题,提出了一种基于骨骼关键点的双标签混合Transformer(dual-label hybrid Transformer,DLHFormer)架构,以实现不同骨架间的直接动作迁移.首先,使用双标签混合Transformer编码器对目标人物骨骼关键点动作序列的动态标签和静态标签进行精确解耦与编码.其次,利用初始姿态预测器预测与动作序列相关联的目标骨骼关键点序列的骨架初始姿态,消除预定义姿态的限制.最后,根据编码后的源人物骨骼的静态标签与目标人物骨骼的动态标签,在框架中实现运动重定向.基于Mixamo构建的跨角色图像数据集进行实验,结果表明:将本研究方法集成到MagicPose后,L1平均损失由 3.69 降至 3.42,PSNR 由 19.92 dB 升至 20.33 dB,SSIM 由 0.996 9 升至 0.997 5,LPIPS 由 0.195 降至0.183,FID由75.5降至68.9;在骨骼关键点重定向任务中,联合空间与时间位置编码时PCK达到0.993,MPJPE降至2.95.研究验证了所提方法能够在跨骨架动作迁移中有效保持人体结构比例,提高关键点定位精度,并改善生成图像的质量.
To address the problem of image generation distortion caused by skeletal proportion differences in human motion transfer,a dual-label hybrid Transformer(DLHFormer)architecture based on skeletal keypoints was proposed to achieve direct motion transfer between different skeletons.First,a DLHFormer encoder was used to accurately decouple and encode the dynamic labels and static labels of the target person's skeletal keypoint motion sequences.Then,an initial pose predictor was used to predict the initial skeleton pose of the target skeletal keypoint sequence associated with the motion sequence,eliminating the limitations of predefined poses.Finally,motion retargeting was realized in the framework based on the encoded static labels of the source person's skeleton and the dynamic labels of the target person's skeleton.Experiments were conducted on a cross-character image dataset constructed based on Mixamo.The results show that after integrating the proposed method into MagicPose,the average L1 loss was reduced from 3.69 to 3.42,PSNR was increased from 19.92 dB to 20.33 dB,SSIM was increased from 0.996 9 to 0.997 5,LPIPS was reduced from 0.195 to 0.183,and FID was reduced from 75.5 to 68.9.In the skeletal keypoint retargeting task,when combining spatial and temporal positional encoding,PCK reached 0.993 and MPJPE was reduced to 2.95.Therefore,the proposed method can effectively maintain human body structure proportions,improve keypoint localization accuracy,and enhance generated image quality in cross-skeleton motion transfer.
刘子壮;丁飞飞
杭州师范大学信息科学与技术学院,浙江杭州杭州师范大学信息科学与技术学院,浙江杭州
信息技术与安全科学
双标签人体动作迁移人体骨骼关键点注意力机制扩散模型
dual-labelhuman motion transferhuman skeletal keypointsattention mechanismdiffusion model
《杭州师范大学学报(自然科学版)》 2026 (4)
367-375,9
浙江省重点研发计划项目(2021c03131)国家自然科学基金项目(62471170)浙江省自然科学基金项目(LQN25F010014,LQN26F020072)
评论