TY - GEN
T1 - Multi-modal Vehicle-Infrastructure Collaborative Perception via Deformable Attention Mechanism
AU - Zhang, Zhenyu
AU - Shi, Junyi
AU - Pang, Haobing
AU - Wang, Mingqian
AU - Zhou, Jianshan
AU - Tian, Daxin
AU - Zheng, Changshui
AU - Liu, Zhiyu
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Multimodal fusion perception techniques have advanced significantly in recent years. However, relying exclusively on ground-based vehicle sensors constrains the perception system to sparse data acquired within the limited field-of-view of terrestrial vehicles, consequently hindering the acquisition of information from alternative perspectives and beyond line-of-sight ranges. This paper presents a novel framework designed to address these issues. We first propose a deformable attention mechanism to adaptively refine feature representations at the local scale. Subsequently, we introduce a context-aware dynamic fusion (CADF) module that dynamically adjusts fusion weights according to environmental context, significantly enhancing both accuracy and robustness. Finally, we present a global-local weighted fusion (GLWF) approach that integrates global channel weighting with local spatial self-attention, effectively optimizing the fusion of perception data from infrastructure and ground vehicle. Experimental results demonstrate that the proposed framework substantially reduces latency, enhances fusion precision, and significantly improves overall perception performance under complex real-world scenarios.
AB - Multimodal fusion perception techniques have advanced significantly in recent years. However, relying exclusively on ground-based vehicle sensors constrains the perception system to sparse data acquired within the limited field-of-view of terrestrial vehicles, consequently hindering the acquisition of information from alternative perspectives and beyond line-of-sight ranges. This paper presents a novel framework designed to address these issues. We first propose a deformable attention mechanism to adaptively refine feature representations at the local scale. Subsequently, we introduce a context-aware dynamic fusion (CADF) module that dynamically adjusts fusion weights according to environmental context, significantly enhancing both accuracy and robustness. Finally, we present a global-local weighted fusion (GLWF) approach that integrates global channel weighting with local spatial self-attention, effectively optimizing the fusion of perception data from infrastructure and ground vehicle. Experimental results demonstrate that the proposed framework substantially reduces latency, enhances fusion precision, and significantly improves overall perception performance under complex real-world scenarios.
KW - attention mechanism
KW - multimodal fusion
KW - perception
KW - vehicle-infrastructure collaboration
UR - https://www.scopus.com/pages/publications/105038330418
U2 - 10.1109/CSIS-IAC65538.2025.11161374
DO - 10.1109/CSIS-IAC65538.2025.11161374
M3 - 会议稿件
AN - SCOPUS:105038330418
T3 - 2025 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2025
SP - 307
EP - 312
BT - 2025 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2025
Y2 - 16 May 2025 through 18 May 2025
ER -