Skip to main navigation Skip to search Skip to main content

Multi-modal Vehicle-Infrastructure Collaborative Perception via Deformable Attention Mechanism

  • Zhenyu Zhang
  • , Junyi Shi
  • , Haobing Pang
  • , Mingqian Wang
  • , Jianshan Zhou*
  • , Daxin Tian
  • , Changshui Zheng
  • , Zhiyu Liu
  • *Corresponding author for this work
  • Beihang University
  • Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Multimodal fusion perception techniques have advanced significantly in recent years. However, relying exclusively on ground-based vehicle sensors constrains the perception system to sparse data acquired within the limited field-of-view of terrestrial vehicles, consequently hindering the acquisition of information from alternative perspectives and beyond line-of-sight ranges. This paper presents a novel framework designed to address these issues. We first propose a deformable attention mechanism to adaptively refine feature representations at the local scale. Subsequently, we introduce a context-aware dynamic fusion (CADF) module that dynamically adjusts fusion weights according to environmental context, significantly enhancing both accuracy and robustness. Finally, we present a global-local weighted fusion (GLWF) approach that integrates global channel weighting with local spatial self-attention, effectively optimizing the fusion of perception data from infrastructure and ground vehicle. Experimental results demonstrate that the proposed framework substantially reduces latency, enhances fusion precision, and significantly improves overall perception performance under complex real-world scenarios.

Original languageEnglish
Title of host publication2025 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages307-312
Number of pages6
ISBN (Electronic)9798331597689
DOIs
StatePublished - 2025
Event2025 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2025 - Shenzhen, China
Duration: 16 May 202518 May 2025

Publication series

Name2025 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2025

Conference

Conference2025 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2025
Country/TerritoryChina
CityShenzhen
Period16/05/2518/05/25

Keywords

  • attention mechanism
  • multimodal fusion
  • perception
  • vehicle-infrastructure collaboration

Fingerprint

Dive into the research topics of 'Multi-modal Vehicle-Infrastructure Collaborative Perception via Deformable Attention Mechanism'. Together they form a unique fingerprint.

Cite this