TY - JOUR
T1 - FaceCLIP
T2 - CLIP-Driven Accurate and Detailed 3D Face Reconstruction from a Single Image
AU - Bao, Yongtang
AU - Zhou, Pengfei
AU - Qi, Liang
AU - Qi, Yue
AU - Li, Haojie
N1 - Publisher Copyright:
© 2024 Tsinghua University Press.
PY - 2026
Y1 - 2026
N2 - In recent years, 3D face reconstruction has become a research hotspot in computer graphics and computer vision. Most current 3DMM-based methods focus on learning displacement maps to recover high-frequency facial details. However, they focus less on learning mid-frequency facial details and introduce displacement maps with noise, decreasing face reconstruction accuracy. Thus, this work presents a novel approach to regressing accurate and detailed 3D face shapes. First, we design a novel feature consistency loss to recover mid-frequency facial details. Specifically, we exploit the powerful CLIP as prior knowledge of faces to extract geometric and semantic features, which helps guide the reconstructed 3D geometric details to match local details in the input image. Furthermore, we propose a parameter refinement module to learn fine-grained features. It helps to obtain accurate model parameters and improve the accuracy of facial reconstruction. Extensive experiments on a FaceScape and a REALY benchmark demonstrate that our method outperforms several state-of-the-art methods in reconstruction accuracy. Furthermore, comprehensive qualitative results show that our approach achieves better visual performance than existing methods.
AB - In recent years, 3D face reconstruction has become a research hotspot in computer graphics and computer vision. Most current 3DMM-based methods focus on learning displacement maps to recover high-frequency facial details. However, they focus less on learning mid-frequency facial details and introduce displacement maps with noise, decreasing face reconstruction accuracy. Thus, this work presents a novel approach to regressing accurate and detailed 3D face shapes. First, we design a novel feature consistency loss to recover mid-frequency facial details. Specifically, we exploit the powerful CLIP as prior knowledge of faces to extract geometric and semantic features, which helps guide the reconstructed 3D geometric details to match local details in the input image. Furthermore, we propose a parameter refinement module to learn fine-grained features. It helps to obtain accurate model parameters and improve the accuracy of facial reconstruction. Extensive experiments on a FaceScape and a REALY benchmark demonstrate that our method outperforms several state-of-the-art methods in reconstruction accuracy. Furthermore, comprehensive qualitative results show that our approach achieves better visual performance than existing methods.
KW - 3D face reconstruction
KW - CLIP
KW - single image
UR - https://www.scopus.com/pages/publications/105029756527
U2 - 10.26599/CVM.2025.9450434
DO - 10.26599/CVM.2025.9450434
M3 - 文章
AN - SCOPUS:105029756527
SN - 2096-0433
VL - 12
SP - 85
EP - 103
JO - Computational Visual Media
JF - Computational Visual Media
IS - 1
ER -