TY - GEN
T1 - MagiaPose
T2 - 2025 China Automation Congress, CAC 2025
AU - Zhou, Hao
AU - Wu, Zhihao
AU - Wang, Jiannan
AU - Zhang, Jian
AU - Sun, Haopeng
AU - Wang, Tian
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - To address the scarcity of objective feedback in training complex somatic skills, such as martial arts, we propose MagiaPose, a video-based analysis framework designed for automated pose recognition and action quality assessment. The framework adopts a multi-task learning paradigm, utilizing a video vision encoder backbone to capture spatiotemporal dependencies from cropped human regions and keyframe features. We evaluate the framework on a martial arts dataset comprising 15 distinct action categories. In our experiments, we compare the performance of fine-tuning a pre-trained ViViT against training lightweight variants (including Attention-based, Mamba-based, and hybrid architectures) from scratch using high-quality, pre-extracted features. Empirical results demonstrate that the Mamba-based variant achieves accuracy nearly comparable to the standard ViViT while significantly reducing model size and inference latency. Notably, the Mamba block exhibits superior accuracy compared to the Attention-based baseline.
AB - To address the scarcity of objective feedback in training complex somatic skills, such as martial arts, we propose MagiaPose, a video-based analysis framework designed for automated pose recognition and action quality assessment. The framework adopts a multi-task learning paradigm, utilizing a video vision encoder backbone to capture spatiotemporal dependencies from cropped human regions and keyframe features. We evaluate the framework on a martial arts dataset comprising 15 distinct action categories. In our experiments, we compare the performance of fine-tuning a pre-trained ViViT against training lightweight variants (including Attention-based, Mamba-based, and hybrid architectures) from scratch using high-quality, pre-extracted features. Empirical results demonstrate that the Mamba-based variant achieves accuracy nearly comparable to the standard ViViT while significantly reducing model size and inference latency. Notably, the Mamba block exhibits superior accuracy compared to the Attention-based baseline.
KW - Mamba
KW - Pose Quality Assessment
KW - Pose Recognition
KW - Video Vision Transformer
UR - https://www.scopus.com/pages/publications/105041163534
U2 - 10.1109/CAC67268.2025.11487683
DO - 10.1109/CAC67268.2025.11487683
M3 - 会议稿件
AN - SCOPUS:105041163534
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 7617
EP - 7622
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 26 September 2025 through 28 September 2025
ER -