跳到主要导航 跳到搜索 跳到主要内容

CLIP-Hand: CLIP-based regressor for hand pose estimation and mesh recovery

  • Feng Zhou
  • , Shuang Ji
  • , Pei Shen
  • , Ju Dai*
  • , Junjun Pan
  • , Yu Kun Lai
  • , Paul L. Rosin
  • *此作品的通讯作者
  • North China University of Technology
  • Peng Cheng Laboratory
  • Cardiff University

科研成果: 期刊稿件文章同行评审

摘要

Despite significant advancements in 3D hand pose estimation, it still faces challenges due to self-occlusion and complex backgrounds. To tackle those issues, we propose a CLIP-based Regressor for Hand Pose Estimation and Mesh Recovery (CLIP-Hand) from a single RGB image. Specifically, we propose an innovative method that combines high-resolution feature aggregation with contrastive language-image pre-trained model (CLIP) to enhance feature representations through language-guided visual prompts. Our approach utilizes a multi-layer Transformer encoder-decoder module to improve the prediction accuracy of hand meshing and joint points. To boost the performance, a predefined 3D joint module and a text dataset are proposed to augment the training data and improve the model’s generalization ability across different scenarios. Extensive experiments on datasets such as FreiHAND, RHD, and Dexter+Object demonstrate the effectiveness of our approach, showing improved performance in terms of accuracy and robustness compared to existing methods. The source code and data will be released once the paper is accepted.

源语言英语
文章编号43
期刊Visual Computer
42
1
DOI
出版状态已出版 - 1月 2026

学术指纹

探究 'CLIP-Hand: CLIP-based regressor for hand pose estimation and mesh recovery' 的科研主题。它们共同构成独一无二的学术指纹。

引用此