跳到主要导航 跳到搜索 跳到主要内容

SIAgent: Spatial Interaction Agent via LLM-Powered Eye-Hand Motion Intent Understanding in VR

  • Zhimin Wang
  • , Chenyu Gu
  • , Feng Lu*
  • *此作品的通讯作者
  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Eye-hand coordinated interaction is becoming a mainstream interaction modality in Virtual Reality (VR) user interfaces. Current paradigms for this multimodal interaction require users to learn predefined gestures and memorize multiple gesture-task associations, which can be summarized as an “Operation-to-Intent” paradigm. This paradigm increases users’ learning costs and has low interaction error tolerance. In this paper, we propose SIAgent, a novel “Intent-to-Operation” framework allowing users to express interaction intents through natural eye-hand motions based on common sense and habits. Our system features two main components: (1) intent recognition that translates spatial interaction data into natural language and infers user intent, and (2) agent-based execution that generates an agent to execute corresponding tasks. This eliminates the need for gesture memorization and accommodates individual motion preferences with high error tolerance. We conduct two user studies across over 60 interaction tasks, comparing our method with two “Operation-to-Intent” techniques. Results show our method achieves higher intent recognition accuracy than gaze + pinch interaction (97.2% versus 93.1%) while reducing arm fatigue and improving usability, and user preference. Another study verifies the function of eye gaze and hand motion channels in intent recognition. Our work offers valuable insights into enhancing VR interaction intelligence through intent-driven design.

源语言英语
页(从-至)6683-6694
页数12
期刊IEEE Transactions on Visualization and Computer Graphics
32
7
DOI
出版状态已出版 - 7月 2026

指纹

探究 'SIAgent: Spatial Interaction Agent via LLM-Powered Eye-Hand Motion Intent Understanding in VR' 的科研主题。它们共同构成独一无二的指纹。

引用此