TY - GEN
T1 - ViMoGen
T2 - International Conference on Extended Reality, ICXR 2025
AU - Wang, Xuehan
AU - Song, Wenfeng
AU - Zhang, Xinyu
AU - Li, Shuai
AU - Wang, Xian’e
AU - Hou, Xia
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Virtual Standard Patient (VSP) is an indispensable tool in medical education, offering crucial experiential learning opportunities. However, current VSP systems struggle to accurately represent patient behaviors and symptoms necessary for real-world diagnostic and assessment tasks. To address this challenge, we introduce ViMoGen, a novel motion generator for VSP driven by a controllable generation pipeline. Specifically, we introduce a conditional control mechanism for our diffusion-based generator. It is guided by dual inputs: instructional text prompts that simulate a physician’s commands, and more critically, expert-defined spatial constraints on specific body joints related to symptoms. This primary contribution allows for the direct encoding of physical limitations, ensuring the generated motions are medically grounded. At the same time, to further enhance the fidelity of the output, we introduce a task-specific loss guidance mechanism, this module refines the initially generated motion by leveraging targeted distance and absolute position losses. This optimization step ensures greater physical plausibility and precision in the final animation. Our experiments demonstrate that by synergistically combining conditional joint control and loss-guided refinement, ViMoGen produces realistic, fine-grained, and medically consistent body motions, making it highly suitable for disease research and medical training scenarios where the interplay of verbal and nonverbal cues is paramount.
AB - Virtual Standard Patient (VSP) is an indispensable tool in medical education, offering crucial experiential learning opportunities. However, current VSP systems struggle to accurately represent patient behaviors and symptoms necessary for real-world diagnostic and assessment tasks. To address this challenge, we introduce ViMoGen, a novel motion generator for VSP driven by a controllable generation pipeline. Specifically, we introduce a conditional control mechanism for our diffusion-based generator. It is guided by dual inputs: instructional text prompts that simulate a physician’s commands, and more critically, expert-defined spatial constraints on specific body joints related to symptoms. This primary contribution allows for the direct encoding of physical limitations, ensuring the generated motions are medically grounded. At the same time, to further enhance the fidelity of the output, we introduce a task-specific loss guidance mechanism, this module refines the initially generated motion by leveraging targeted distance and absolute position losses. This optimization step ensures greater physical plausibility and precision in the final animation. Our experiments demonstrate that by synergistically combining conditional joint control and loss-guided refinement, ViMoGen produces realistic, fine-grained, and medically consistent body motions, making it highly suitable for disease research and medical training scenarios where the interplay of verbal and nonverbal cues is paramount.
KW - Controllable Human Motion Synthesis
KW - Diffusion Model
KW - Virtual Standard Patient
UR - https://www.scopus.com/pages/publications/105041985741
U2 - 10.1007/978-981-95-7195-6_1
DO - 10.1007/978-981-95-7195-6_1
M3 - 会议稿件
AN - SCOPUS:105041985741
SN - 9789819571949
T3 - Lecture Notes in Computer Science
SP - 1
EP - 17
BT - Extended Reality - International Conference, ICXR 2025, Proceedings
A2 - Hinkenjan, Andre
A2 - Liang, Hai-Ning
A2 - Qin, Xueying
A2 - Zhang, Song-Hai
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 1 November 2025 through 2 November 2025
ER -