TY - JOUR
T1 - LM2CNet
T2 - Enhancing monocular 3D visual grounding with language guided multi-modality coupling network
AU - Li, Meng
AU - Zhao, Qi
AU - Lyu, Shuchang
AU - Jiang, Jun
AU - Zou, Longhao
AU - Cheng, Guangliang
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/7
Y1 - 2026/7
N2 - Monocular 3D object detection, an endeavor to predict 3D bounding boxes from single images, has seen substantial advancements with the introduction of Mono3DVG, which integrates language descriptions for more precise object localization in 3D scenes. Despite its initial successes, challenges persist with the limitations of existing datasets and vanilla architectural designs. Addressing these, we introduce a language guided Multi-Modality Coupling Network (LM2CNet) and a new large-scale benchmark dataset, Mono3DRefer-nuScenes, which utilizes Large Language Model (LLM) to generate diverse language descriptions for enhancing the evaluation of Mono3D visual grounding tasks. Our dataset demonstrates a significant improvement over previous datasets by expanding the number of objects, instances, and distance ranges included. Architecturally, LM2CNet innovatively employs a two-stage cross-modality coupling module that synergizes visual and depth features through language embedding, serving as an intermediate modality. This design optimizes the integration of language and depth information for guiding visual features. Our extensive experimental results confirm the effectiveness of LM2CNet, which achieves state-of-the-art performance with an average accuracy of 49.50% and 27.34% on Acc0.25 and Acc0.5 metrics in Mono3DRefer-nuScenes, respectively. Our code and dataset are available at https://github.com/cv516Buaa/LM2CNet
AB - Monocular 3D object detection, an endeavor to predict 3D bounding boxes from single images, has seen substantial advancements with the introduction of Mono3DVG, which integrates language descriptions for more precise object localization in 3D scenes. Despite its initial successes, challenges persist with the limitations of existing datasets and vanilla architectural designs. Addressing these, we introduce a language guided Multi-Modality Coupling Network (LM2CNet) and a new large-scale benchmark dataset, Mono3DRefer-nuScenes, which utilizes Large Language Model (LLM) to generate diverse language descriptions for enhancing the evaluation of Mono3D visual grounding tasks. Our dataset demonstrates a significant improvement over previous datasets by expanding the number of objects, instances, and distance ranges included. Architecturally, LM2CNet innovatively employs a two-stage cross-modality coupling module that synergizes visual and depth features through language embedding, serving as an intermediate modality. This design optimizes the integration of language and depth information for guiding visual features. Our extensive experimental results confirm the effectiveness of LM2CNet, which achieves state-of-the-art performance with an average accuracy of 49.50% and 27.34% on Acc0.25 and Acc0.5 metrics in Mono3DRefer-nuScenes, respectively. Our code and dataset are available at https://github.com/cv516Buaa/LM2CNet
KW - 3D visual grounding
KW - Cross modality coupling
KW - Language description
KW - Large language model
KW - Monocular 3D object detection
UR - https://www.scopus.com/pages/publications/105038555947
U2 - 10.1016/j.patrec.2026.03.023
DO - 10.1016/j.patrec.2026.03.023
M3 - 文章
AN - SCOPUS:105038555947
SN - 0167-8655
VL - 205
SP - 109
EP - 115
JO - Pattern Recognition Letters
JF - Pattern Recognition Letters
ER -