Skip to main navigation Skip to search Skip to main content

LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network

  • Meng Li
  • , Qi Zhao
  • , Shuchang Lyu
  • , Jun Jiang*
  • , Longhao Zou
  • , Guangliang Cheng
  • *Corresponding author for this work
  • Beihang University
  • Pengcheng Laboratory
  • University of Liverpool

Research output: Contribution to journalArticlepeer-review

Abstract

Monocular 3D object detection, an endeavor to predict 3D bounding boxes from single images, has seen substantial advancements with the introduction of Mono3DVG, which integrates language descriptions for more precise object localization in 3D scenes. Despite its initial successes, challenges persist with the limitations of existing datasets and vanilla architectural designs. Addressing these, we introduce a language guided Multi-Modality Coupling Network (LM2CNet) and a new large-scale benchmark dataset, Mono3DRefer-nuScenes, which utilizes Large Language Model (LLM) to generate diverse language descriptions for enhancing the evaluation of Mono3D visual grounding tasks. Our dataset demonstrates a significant improvement over previous datasets by expanding the number of objects, instances, and distance ranges included. Architecturally, LM2CNet innovatively employs a two-stage cross-modality coupling module that synergizes visual and depth features through language embedding, serving as an intermediate modality. This design optimizes the integration of language and depth information for guiding visual features. Our extensive experimental results confirm the effectiveness of LM2CNet, which achieves state-of-the-art performance with an average accuracy of 49.50% and 27.34% on Acc0.25 and Acc0.5 metrics in Mono3DRefer-nuScenes, respectively. Our code and dataset are available at https://github.com/cv516Buaa/LM2CNet

Original languageEnglish
Pages (from-to)109-115
Number of pages7
JournalPattern Recognition Letters
Volume205
DOIs
StatePublished - Jul 2026

Keywords

  • 3D visual grounding
  • Cross modality coupling
  • Language description
  • Large language model
  • Monocular 3D object detection

Fingerprint

Dive into the research topics of 'LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network'. Together they form a unique fingerprint.

Cite this