首页|期刊导航|郑州大学学报(工学版)|基于扩散模型和交叉注意力机制的骨骼点动作识别方法

基于扩散模型和交叉注意力机制的骨骼点动作识别方法OA

Diffusion Method and Cross-attention Mechanisms for Skeleton-based Action Recognition Method

中文摘要英文摘要

针对人体动作识别中骨骼序列因遮挡或关节点缺失导致的动作信息不完整,以及在标注样本有限情况下模型泛化能力不足等问题,提出了一种结合扩散模型和交叉注意力机制的骨骼点动作识别方法 DCMAE.在自监督学习框架下,采用时空掩蔽策略,通过扩散模型在去噪过程中学习动作序列的全局分布特性,提升模型在数据缺损情况下的分类准确率;在解码阶段通过交叉注意力机制引入编码器特征,实现时空维度的信息交互与引导,从而增强模型在少标签条件下的泛化能力.实验在 NTU RGB+D 60 和 NTU RGB+D 120 数据集上进行,结果表明:所提方法在数据缺损情况下和少标签条件下的识别准确率较 SkeletonMAE 最高分别提升了 14.9 百分点和 3.0 百分点.所提方法能够有效增强骨骼动作识别模型对缺损数据和少标签数据的鲁棒性,为自监督动作识别提供了新思路.

To address the problems of incomplete motion information caused by occlusion or missing joints in skele-ton-based action recognition,as well as the limited generalization ability of models with few-label conditions,a skeleton-based action recognition method DCMAE was proposed,which integrated a diffusion model with a cross-at-tention mechanism.Within a self-supervised learning framework,a spatio-temporal masking strategy was adopted,where the diffusion model learned the global distribution characteristics of motion sequences during the denoising process to improve classification accuracy under data-missing conditions.In the decoding stage,the cross-attention mechanism introduced encoder features to achieve spatio-temporal information interaction and guidance,thereby en-hancing the model's generalization ability in few-label conditions.Experiments conducted on the NTU RGB+D 60 and NTU RGB+D 120 datasets showed that the proposed method could achieve accuracy improvements of up to 14.9 percentage points and 3.0 percentage points,respectively,over SkeletonMAE with data-missing conditions and few-label conditions.The proposed method effectively enhanced the robustness of skeleton-based action recog-nition models to data-missing and few-label data,providing a new perspective for self-supervised action recognition research.

陈恩庆;李佳惠;郭新

郑州大学 电气与信息工程学院,河南 郑州 450001郑州大学 电气与信息工程学院,河南 郑州 450001郑州大学 电气与信息工程学院,河南 郑州 450001

信息技术与安全科学

骨骼点动作识别自监督学习掩蔽重建扩散模型交叉注意力机制

skeleton-based action recognitionself-supervised learningmasked reconstructiondiffusion modelcross-attention mechanism

《郑州大学学报(工学版)》 2026 (4)

1-8,8

国家自然科学基金资助项目(62101503)河南省科技攻关计划项目(242102211017)

10.13705/j.issn.1671-6833.2026.04.011

评论