TY - JOUR
T1 - SwinFaceNet
T2 - Segmentation-guided transformer for UV-based 3D face reconstruction
AU - Zhao, Yaopu
AU - Gong, Guanghong
AU - Li, Ni
N1 - Publisher Copyright:
© 2026 Published by Elsevier Ltd.
PY - 2026/6
Y1 - 2026/6
N2 - Single-image 3D face reconstruction is essential for a wide range of graphics applications, including facial animation, virtual human creation, and immersive augmented and virtual reality (AR/VR) experiences. While position map regression has emerged as an effective paradigm, existing CNN-based encoder–decoder architectures struggle to model long-range geometric dependencies, and their feature responses often spread into non-facial regions, compromising both accuracy and interpretability. To address these issues, we propose SwinFaceNet, a segmentation-guided Transformer framework for monocular 3D face reconstruction based on position maps. Our model integrates a hierarchical Swin Transformer encoder with a multi-scale Transformer decoder, enabling cross-scale feature fusion through self- and cross-attention. Furthermore, we introduce a face-aware attention supervision mechanism that leverages face-region masks to regularize the spatial distribution of attention, encouraging the model to focus on geometrically meaningful facial regions. Experiments on the AFLW2000-3D benchmark show that SwinFaceNet achieves strong performance in both 2D landmark alignment and 3D face reconstruction, while providing visually interpretable attention maps that enhance model transparency and diagnostics. Our work demonstrates the potential of Transformer-based architectures with explicit structural guidance for dense facial geometry reconstruction in UV space, while highlighting a promising direction toward more interpretable 3D face modeling.
AB - Single-image 3D face reconstruction is essential for a wide range of graphics applications, including facial animation, virtual human creation, and immersive augmented and virtual reality (AR/VR) experiences. While position map regression has emerged as an effective paradigm, existing CNN-based encoder–decoder architectures struggle to model long-range geometric dependencies, and their feature responses often spread into non-facial regions, compromising both accuracy and interpretability. To address these issues, we propose SwinFaceNet, a segmentation-guided Transformer framework for monocular 3D face reconstruction based on position maps. Our model integrates a hierarchical Swin Transformer encoder with a multi-scale Transformer decoder, enabling cross-scale feature fusion through self- and cross-attention. Furthermore, we introduce a face-aware attention supervision mechanism that leverages face-region masks to regularize the spatial distribution of attention, encouraging the model to focus on geometrically meaningful facial regions. Experiments on the AFLW2000-3D benchmark show that SwinFaceNet achieves strong performance in both 2D landmark alignment and 3D face reconstruction, while providing visually interpretable attention maps that enhance model transparency and diagnostics. Our work demonstrates the potential of Transformer-based architectures with explicit structural guidance for dense facial geometry reconstruction in UV space, while highlighting a promising direction toward more interpretable 3D face modeling.
KW - 3D face reconstruction
KW - Attention supervision
KW - Swin Transformer
KW - UV position map
UR - https://www.scopus.com/pages/publications/105039263165
U2 - 10.1016/j.cag.2026.104626
DO - 10.1016/j.cag.2026.104626
M3 - 文章
AN - SCOPUS:105039263165
SN - 0097-8493
VL - 137
JO - Computers and Graphics
JF - Computers and Graphics
M1 - 104626
ER -