Skip to main navigation Skip to search Skip to main content

EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion

  • Xin Zhao
  • , Ju Dai*
  • , Feng Zhou
  • , Haofei Wang
  • , Zhen Song
  • , Aimin Hao
  • , Hong Qin
  • , Yang Gao*
  • *Corresponding author for this work
  • Beihang University
  • Peng Cheng Laboratory
  • North China University of Technology
  • Stony Brook University

Research output: Contribution to journalArticlepeer-review

Abstract

Speech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies.

Original languageEnglish
Pages (from-to)7631-7644
Number of pages14
JournalIEEE Transactions on Visualization and Computer Graphics
Volume32
Issue number8
DOIs
StateAccepted/In press - 2026

Keywords

  • 3D facial animation
  • fine-grained edit
  • speech driven

Fingerprint

Dive into the research topics of 'EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion'. Together they form a unique fingerprint.

Cite this