跳到主要导航 跳到搜索 跳到主要内容

LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network

  • Meng Li
  • , Qi Zhao
  • , Shuchang Lyu
  • , Jun Jiang*
  • , Longhao Zou
  • , Guangliang Cheng
  • *此作品的通讯作者
  • Beihang University
  • Pengcheng Laboratory
  • University of Liverpool

科研成果: 期刊稿件文章同行评审

摘要

Monocular 3D object detection, an endeavor to predict 3D bounding boxes from single images, has seen substantial advancements with the introduction of Mono3DVG, which integrates language descriptions for more precise object localization in 3D scenes. Despite its initial successes, challenges persist with the limitations of existing datasets and vanilla architectural designs. Addressing these, we introduce a language guided Multi-Modality Coupling Network (LM2CNet) and a new large-scale benchmark dataset, Mono3DRefer-nuScenes, which utilizes Large Language Model (LLM) to generate diverse language descriptions for enhancing the evaluation of Mono3D visual grounding tasks. Our dataset demonstrates a significant improvement over previous datasets by expanding the number of objects, instances, and distance ranges included. Architecturally, LM2CNet innovatively employs a two-stage cross-modality coupling module that synergizes visual and depth features through language embedding, serving as an intermediate modality. This design optimizes the integration of language and depth information for guiding visual features. Our extensive experimental results confirm the effectiveness of LM2CNet, which achieves state-of-the-art performance with an average accuracy of 49.50% and 27.34% on Acc0.25 and Acc0.5 metrics in Mono3DRefer-nuScenes, respectively. Our code and dataset are available at https://github.com/cv516Buaa/LM2CNet

源语言英语
页(从-至)109-115
页数7
期刊Pattern Recognition Letters
205
DOI
出版状态已出版 - 7月 2026

学术指纹

探究 'LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network' 的科研主题。它们共同构成独一无二的学术指纹。

引用此