Skip to main navigation Skip to search Skip to main content

Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes

  • Yin Wang*
  • , Zhiying Leng
  • , Haitian Liu
  • , Frederick W.B. Li
  • , Mu Li
  • , Xiaohui Liang*
  • *Corresponding author for this work
  • Beihang University
  • Durham University
  • Zhongguancun Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

Scenes are continuously undergoing dynamic changes in the real world. However, existing human-scene interaction generation methods typically treat the scene as static, which deviates from reality. Inspired by world models, we introduce Dyn-HSI, the first cognitive architecture for dynamic human-scene interaction, which endows virtual humans with three humanoid components. (1) Vision (human eyes): we equip the virtual human with a Dynamic Scene-Aware Navigation, which continuously perceives changes in the surrounding environment and adaptively predicts the next waypoint. (2) Memory (human brain): we equip the virtual human with a Hierarchical Experience Memory, which stores and updates experiential data accumulated during training. This allows the model to leverage prior knowledge during inference for context-aware motion priming, thereby enhancing both motion quality and generalization. (3) Control (human body): we equip the virtual human with Human-Scene Interaction Diffusion Model, which generates high-fidelity interaction motions conditioned on multimodal inputs. To evaluate performance in dynamic scenes, we extend the existing static human-scene interaction datasets to construct a dynamic benchmark, Dyn-Scenes. We conduct extensive qualitative and quantitative experiments to validate Dyn-HSI, showing that our method consistently outperforms existing approaches and generates high-quality human-scene interaction motions in both static and dynamic settings.

Original languageEnglish
Pages (from-to)3062-3072
Number of pages11
JournalIEEE Transactions on Visualization and Computer Graphics
Volume32
Issue number5
DOIs
StatePublished - 1 May 2026

Keywords

  • Diffusion Model
  • Dynamic Scene
  • Human-Scene Interaction
  • Text-Driven Motion Generation
  • World Model

Fingerprint

Dive into the research topics of 'Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes'. Together they form a unique fingerprint.

Cite this