TY - JOUR
T1 - Multimodal 3D object detection for autonomous driving under vision-language supervision
T2 - a contrastive-learning perspective
AU - Lin, Chunmian
AU - Zhang, Wenze
AU - Chen, Yanyan
AU - Yang, Lei
AU - Jiang, Han
AU - Tian, Daxin
AU - Duan, Xuting
AU - Zhou, Jianshan
AU - Cao, Dongpu
N1 - Publisher Copyright:
© Science China Press 2026.
PY - 2026/5
Y1 - 2026/5
N2 - Multimodal large language models (MLLMs) have been well acknowledged as the generalist across a broad spectrum of vision-language understanding tasks. Despite notable advancements, their potential for autonomous-driving perception remains largely underexplored. In response, we conduct an in-depth investigation of image-text-point interaction and propose a versatile paradigm of vision-language supervision (VLS) for 3D object detection, where multi-sensory proposals are primarily refined with meticulously-designed text-referred expression, and multimodal correspondences are further incorporated in a contrastive-learning manner. Moreover, VLS holds great advantages. (1) No complicated engineering. It could be seamlessly integrated into a camera-LiDAR 3D detector without troublesome hand-crafted engineering. (2) No extra computation. It provides auxiliary guidance only during training. (3) No additional data. It derives multimodal pairs from ground-truth label instead of a laborious annotation pipeline. Empirical study on publicly available KITTI and nuScenes benchmarks demonstrates the state-of-the-art detection performance against a wide span of counterparts, suggesting its effectiveness and advancement. We hope this work could pave a substantial path towards multimodal feature fusion and object detection for autonomous driving.
AB - Multimodal large language models (MLLMs) have been well acknowledged as the generalist across a broad spectrum of vision-language understanding tasks. Despite notable advancements, their potential for autonomous-driving perception remains largely underexplored. In response, we conduct an in-depth investigation of image-text-point interaction and propose a versatile paradigm of vision-language supervision (VLS) for 3D object detection, where multi-sensory proposals are primarily refined with meticulously-designed text-referred expression, and multimodal correspondences are further incorporated in a contrastive-learning manner. Moreover, VLS holds great advantages. (1) No complicated engineering. It could be seamlessly integrated into a camera-LiDAR 3D detector without troublesome hand-crafted engineering. (2) No extra computation. It provides auxiliary guidance only during training. (3) No additional data. It derives multimodal pairs from ground-truth label instead of a laborious annotation pipeline. Empirical study on publicly available KITTI and nuScenes benchmarks demonstrates the state-of-the-art detection performance against a wide span of counterparts, suggesting its effectiveness and advancement. We hope this work could pave a substantial path towards multimodal feature fusion and object detection for autonomous driving.
KW - adapter
KW - autonomous driving
KW - contrastive learning
KW - multimodal 3D detection
KW - vision-language model
UR - https://www.scopus.com/pages/publications/105037456342
U2 - 10.1007/s11432-025-4853-5
DO - 10.1007/s11432-025-4853-5
M3 - 文章
AN - SCOPUS:105037456342
SN - 1674-733X
VL - 69
JO - Science China Information Sciences
JF - Science China Information Sciences
IS - 5
M1 - 150106
ER -