TY - JOUR
T1 - EmoPoseFace
T2 - Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion
AU - Zhao, Xin
AU - Dai, Ju
AU - Zhou, Feng
AU - Wang, Haofei
AU - Song, Zhen
AU - Hao, Aimin
AU - Qin, Hong
AU - Gao, Yang
N1 - Publisher Copyright:
© 1995-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Speech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies.
AB - Speech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies.
KW - 3D facial animation
KW - fine-grained edit
KW - speech driven
UR - https://www.scopus.com/pages/publications/105041531435
U2 - 10.1109/TVCG.2026.3702133
DO - 10.1109/TVCG.2026.3702133
M3 - 文章
AN - SCOPUS:105041531435
SN - 1077-2626
VL - 32
SP - 7631
EP - 7644
JO - IEEE Transactions on Visualization and Computer Graphics
JF - IEEE Transactions on Visualization and Computer Graphics
IS - 8
ER -