Skip to main navigation Skip to search Skip to main content

FCMD: Fine-Grained Text-Driven Cohesive Motion Generation With Diffusion Model

  • Shuai Li
  • , Siqi Wang
  • , Xinyu Zhang
  • , Wenfeng Song*
  • , Aimin Hao
  • *Corresponding author for this work
  • Beihang University
  • Beijing Information Science & Technology University

Research output: Contribution to journalArticlepeer-review

Abstract

Generating continuous and expressive human motion from textual descriptions is a critical challenge in applications such as gaming and filmmaking. Existing methods often struggle to maintain global coherence, realistic frame continuity, and smooth transitions. To address these limitations, we propose FCMD, a novel diffusion-based model for generating cohesive motion sequences from fine-grained textual descriptions. FCMD introduces three key innovations: (1) Fine-grained Text Fusion, which integrates detailed textual cues with transitional narratives to enhance semantic consistency; (2) History Motion Guidance, ensuring motion accuracy and consistency across consecutive frames; and (3) Smooth Stitching Sampling, which leverages preceding and current motion information to achieve seamless transitions. Additionally, FCMD employs a large language model (LLM) to refine motion datasets by extracting fine-grained textual descriptions. Extensive experiments demonstrate that FCMD outperforms state-of-the-art methods in generating coherent, natural, and highly controllable motion sequences.

Original languageEnglish
JournalIEEE Transactions on Visualization and Computer Graphics
DOIs
StateAccepted/In press - 2026

Keywords

  • Diffusion model
  • human motion synthesis
  • text-driven generation

Fingerprint

Dive into the research topics of 'FCMD: Fine-Grained Text-Driven Cohesive Motion Generation With Diffusion Model'. Together they form a unique fingerprint.

Cite this