Skip to main navigation Skip to search Skip to main content

3DAxisPrompt: Promoting the 3D grounding and reasoning in GPT-4o

  • Dingning Liu
  • , Cheng Wang
  • , Peng Gao
  • , Renrui Zhang
  • , Xinzhu Ma*
  • , Yuan Meng
  • , Zhihui Wang
  • *Corresponding author for this work
  • Shanghai Artificial Intelligence Laboratory
  • Dalian University of Technology
  • Wuhan University
  • Chinese University of Hong Kong
  • Tsinghua University

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual understanding, while the capability of MLLMs to operate effectively in 3D vision remains an ongoing area of exploration. In this paper, we introduce a novel visual prompting method, called 3DAxisPrompt, to elicit the 3D understanding capabilities of MLLMs in real-world scenes. More specifically, our method leverages the 3D coordinate axis and masks generated from the Segment Anything Model (SAM) to provide explicit geometric priors to MLLMs and then extend their impressive 2D grounding/reasoning ability to real-world 3D scenarios. Besides, we first provide a thorough investigation of the potential visual prompting formats and conclude our findings to reveal the potential and limits of 3D understanding capabilities in GPT-4o, as a representative of MLLMs. Finally, we build evaluation environments with four datasets, i.e. ScanRefer, ScanNet, FMB, and nuScene datasets, covering various 3D tasks. Based on this, we conduct extensive quantitative and qualitative experiments, which demonstrate the effectiveness of the proposed method. Overall, our study reveals that MLLMs, with the help of 3DAxisPrompt, can effectively perceive an object's 3D position in real-world scenarios. Nevertheless, a single prompt engineering approach does not consistently achieve the best outcomes for all 3D tasks. This study highlights the feasibility of leveraging MLLMs for 3D vision grounding/reasoning with prompt engineering techniques.

Original languageEnglish
Article number130072
JournalNeurocomputing
Volume637
DOIs
StatePublished - 7 Jul 2025
Externally publishedYes

Keywords

  • 3D grounding
  • GPT-4o
  • MLLMs
  • Visual prompt

Fingerprint

Dive into the research topics of '3DAxisPrompt: Promoting the 3D grounding and reasoning in GPT-4o'. Together they form a unique fingerprint.

Cite this