基于通道自注意和时空并行感知的体育动作质量评估网络

    Sports action quality assessment network based on tube self-attention and spatiotemporal parallel perception

    • 摘要: 体育动作质量评估需要在动作类型识别的基础上对动作的完整性、流畅性、难易程度等给出质量评分,挑战性极大。为准确提取动作序列中的关键信息、高效评估动作质量,设计了一种基于通道自注意和时空并行感知的网络结构。首先,使用单目标跟踪与通道自注意模块提取动作信息,排除背景信息干扰;其次,以Uniformer模型为基础,在局部特征提取模块引入空间多头自注意力机制,将局部和全局注意力时空串行感知网络改为时空并行感知网络,以提升计算效率;最后,通过多阶段融合模块将局部和全局特征融合以增强动作特征,并扩展多层感知模块,输出动作识别和动作质量评估结果。在UCF101与HMDB51数据集上进行动作识别,UCF101的动作识别在Top-1与Top-5上准确率分别达到85.2%与96.7%,运算成本FLOPs明显下降,HMDB51的动作识别在Top-1上准确率达到77.8%;在AQA-7与MTL-AQA数据集上进行动作质量评估实验,AQA-7的平均斯皮尔曼等级相关系数达到0.808 7,MTL-AQA的斯皮尔曼等级相关系数达到0.948 1。实验结果表明,该模型不仅有较高的动作识别率,也能准确高效地完成动作质量评估任务,并且在多个数据集上性能表现稳定,表明了模型的泛化性较好。

       

      Abstract: The quality assessment of sports actions requires providing quality ratings for the completeness, fluency, and difficulty of actions based on the recognition of action types, which is extremely challenging. A network structure based on tube self-attention and spatiotemporal parallel perception was designed to accurately extract key information from action sequences and efficiently assess action quality. Firstly, action information was extracted and background interference was eliminated using a single target tracking and tube self-attention module. Secondly, based on the Uniformer model, a spatial multi-head self-attention mechanism was introduced in the local feature extraction module, and the local and global attention spatiotemporal serial perception network was changed to a spatiotemporal parallel perception network to improve computational efficiency. Finally, the local and global features were fused through a multi-stage fusion module to enhance action representation, and the multi-layer perceptron module was extended to output action recognition and action quality assessment results. Action recognition experiments were conducted on the UCF101 and HMDB51 dataset, respectively. The action recognition Top-1 and Top-5 accuracy of the UCF101 dataset reached 85.2% and 96.7%, respectively, and the computational cost of FLOPs significantly decreased, the action recognition accuracy Top-1 of the HMDB51 dataset reached 77.8%. Action quality assessment experiments were conducted on the AQA-7 and MTL-AQA dataset, respectively, the average Spearman’s rank correlation coefficient for the assessment of the AQA-7 dataset reached 0.808 7, the Spearman’s rank correlation coefficient for the assessment of the MTL-AQA dataset reached 0.948 1. The experimental results show that the model not only has a high action recognition rate, but also can accurately and efficiently complete the task of action quality assessment.

       

    /

    返回文章
    返回