基于自适应通道特征交互融合图卷积网络的骨骼行为识别OA
Adaptive Channel Feature Interactive Fusion Network for Skeleton-based Action Recognition
本研究提出了一种基于图卷积网络(Graph Convolutional Networks,GCN)的骨骼行为识别方法,针对传统时空图卷积框架在处理时空特征时存在的统一化处理及忽视通道间交互的问题,提出了一种有效的改进方案.模型通过融合多种拓扑矩阵,增强了空间信息的互补表达,同时引入了通道交互注意力模块,通过捕捉时空维度中的帧间动态信息和人体结构特征,建模不同通道间的交互关系,提升特征表达能力.此外,模型设计的时间自适应特征融合模块(Temporal Adaptive Feature Fusion,TAF)通过自适应选择不同网络层中的扩张率和卷积核大小,解决了上下文聚合和初始特征集成的问题.TAF模块分别关注初始特征和时间维度的信息,进行有效的特征融合,成功整合了初始特征与高维时间特征,从而显著提高了时空特征提取的能力.在NW-UCLA数据集上,所提出的方法相比基准模型CTR-GCN(Channel-wise Topology Refinement Graph Convolution Network)提升了2.1%的识别精度,相较最新方法Info-GCN提高了0.7%.在NTU RGB+D 120和NTU RGB+D数据集的不同划分方式下,模型分别比基础模型识别准确率提高了0.7%、0.8%及0.5%、0.6%,并在各项评价指标上均超过了现有最新方法.实验结果表明,所提出的模型在时空特征提取和骨骼行为识别任务中均表现出显著的性能优势.
This study proposed a novel skeleton-based action recognition method utilizing Graph Convolutional Networks(GCN),which addressed the limitations of conventional spatiotemporal graph convolution frameworks that uniformly process spatiotempo-ral features while neglecting inter-channel interactions.Specifically,the proposed model enhanced the complementary representation of spatial information through the fusion of multiple topological matrices coupled with the introduction of a Channel Interaction At-tention(CIA)module.The CIA module was designed to capture dynamic frame-level information and human structural features across spatiotemporal dimensions,effectively modeling inter-channel relationships and thereby improving skeletal data representa-tion.Furthermore,a Temporal Adaptive Feature Fusion(TAF)module was incorporated to adaptively select varying dilation rates and kernel sizes across network layers.This module replaced traditional residual connections between initial features and temporal module outputs,effectively addressing context aggregation and initial feature integration challenges.The TAF module separately pro-cessed initial features and temporal information,enabling efficient feature fusion and successful integration of initial features with high-dimensional temporal features,which significantly enhanced spatiotemporal feature extraction.Experimental results demon-strated that on the NW-UCLA dataset,the proposed method achieved 2.1%higher recognition accuracy than the baseline model CTR-GCN(Channel-wise Topology Refinement Graph Convolution Network)and 0.7%improvement over state-of-the-art methods Info-GCN.For the NTU RGB+D 120 and NTU RGB+D datasets under different splits,the model showed consistent performance gains of 0.7%,0.8%and 0.5%,0.6%,respectively,surpassing all existing methods across evaluation metrics.These results con-firmed the model's superior performance in both spatiotemporal feature extraction and skeleton-based action recognition tasks.
施宇航;陈琳琳;郭峰;何强
北京建筑大学 理学院,北京 102616北京建筑大学 理学院,北京 102616||北京建筑大学 大数据建模理论与技术研究所,北京 102616奇安信科技集团,北京 100044北京建筑大学 理学院,北京 102616||北京建筑大学 大数据建模理论与技术研究所,北京 102616
信息技术与安全科学
骨骼行为识别通道注意力机制时空特征融合通道交互
skeleton-based action recognitionchannel attention mechanismspatiotemporal feature fusionchannel interaction
《山西大学学报(自然科学版)》 2026 (2)
220-231,12
国家自然科学基金(12301581)北京市自然科学基金(4252033)北京市教育委员会科学研究计划项目(KM202210016002)北京建筑大学基本科研业务费资助(X25039)北京建筑大学硕士研究生创新项目(PG2025172)
评论