跳到主要导航 跳到搜索 跳到主要内容

AsynFusion: Towards Asynchronous Latent Consistency Models for Decoupled Whole-Body Audio-Driven Avatars

  • Tianbao Zhang
  • , Jian Zhao
  • , Yuer Li
  • , Zheng Zhu
  • , Ping Hu
  • , Zhaoxin Fan*
  • , Wenjun Wu
  • , Xuelong Li*
  • *此作品的通讯作者
  • China Telecommunications
  • GigaAI
  • Xinjiang University
  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital entertainment, and remote communication. Existing approaches often generate audio-driven facial expressions and gestures independently, which introduces a significant limitation: the lack of seamless coordination between facial and gestural elements, resulting in less natural and cohesive animations. To address this limitation, we propose AsynFusion, a novel framework that leverages diffusion transformers to achieve harmonious expression and gesture synthesis. The proposed method is built upon a dual-branch DiT architecture, which enables the parallel generation of facial expressions and gestures. Within the model, we introduce a Cooperative Synchronization Module to facilitate bidirectional feature interaction between the two modalities, and an Asynchronous LCM Sampling strategy to reduce computational overhead while maintaining high-quality outputs. Extensive experiments demonstrate that AsynFusion achieves state-of-the-art performance in generating real-time, synchronized whole-body animations, consistently outperforming existing methods in both quantitative and qualitative evaluations.

源语言英语
主期刊名Pattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings
编辑Josef Kittler, Hongkai Xiong, Weiyao Lin, Jian Yang, Xilin Chen, Jiwen Lu, Jingyi Yu, Weishi Zheng
出版商Springer Science and Business Media Deutschland GmbH
519-535
页数17
ISBN(印刷版)9789819556755
DOI
出版状态已出版 - 2026
活动8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025 - Shanghai, 中国
期限: 15 10月 202518 10月 2025

出版系列

姓名Lecture Notes in Computer Science
16278 LNCS
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025
国家/地区中国
Shanghai
时期15/10/2518/10/25

指纹

探究 'AsynFusion: Towards Asynchronous Latent Consistency Models for Decoupled Whole-Body Audio-Driven Avatars' 的科研主题。它们共同构成独一无二的指纹。

引用此