Abstract
Monocular 3D object detection, an endeavor to predict 3D bounding boxes from single images, has seen substantial advancements with the introduction of Mono3DVG, which integrates language descriptions for more precise object localization in 3D scenes. Despite its initial successes, challenges persist with the limitations of existing datasets and vanilla architectural designs. Addressing these, we introduce a language guided Multi-Modality Coupling Network (LM2CNet) and a new large-scale benchmark dataset, Mono3DRefer-nuScenes, which utilizes Large Language Model (LLM) to generate diverse language descriptions for enhancing the evaluation of Mono3D visual grounding tasks. Our dataset demonstrates a significant improvement over previous datasets by expanding the number of objects, instances, and distance ranges included. Architecturally, LM2CNet innovatively employs a two-stage cross-modality coupling module that synergizes visual and depth features through language embedding, serving as an intermediate modality. This design optimizes the integration of language and depth information for guiding visual features. Our extensive experimental results confirm the effectiveness of LM2CNet, which achieves state-of-the-art performance with an average accuracy of 49.50% and 27.34% on Acc0.25 and Acc0.5 metrics in Mono3DRefer-nuScenes, respectively. Our code and dataset are available at https://github.com/cv516Buaa/LM2CNet
| Original language | English |
|---|---|
| Pages (from-to) | 109-115 |
| Number of pages | 7 |
| Journal | Pattern Recognition Letters |
| Volume | 205 |
| DOIs | |
| State | Published - Jul 2026 |
Keywords
- 3D visual grounding
- Cross modality coupling
- Language description
- Large language model
- Monocular 3D object detection
Fingerprint
Dive into the research topics of 'LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver