Kinematic Enhanced Hypergraph Convolutional Network for Skeleton-based Human Action Recognition with LLM Training Guides
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755538 ↗
摘要
Skeleton-based human action recognition has wide applications in video understanding and virtual reality. However, most existing methods focus excessively on spatial location and global movement, while underrepresenting subtle and local actions. To address the limitation, we innovatively propose a Kinematic Enhanced Hypergraph Convolutional Network(KEHCN) with LLM training guides. The network mainly consists of LLM Training Guides(LTG), Kinematic Hypergraph Convolution(KHC), and Kinematic Gating Module(KGM). Specifically, we use the hypergraph convolutional network to extract high-order correlated human skeleton features, the KHC to encode the kinematic features and the LTG to provide a pre-trained large language model to generate text and kinematic description features during the training phase. Based on the Mixture of Experts (MoE) framework, we simplify the gating network by introducing a kinematic feature threshold, thereby constructing a dual-branch global and local motion expert network (KGM). We integrated kinematic features into KHC, LTG and KGM to seek improvements from three perspectives, all of which have enhanced the performance. The experiments on three benchmark datasets(NTU RGB+D, NTU-RGB+D 120 and NW-UCLA), demonstrate the state-of-the-art performance compared to current open-source methods.