跳到主要导航 跳到搜索 跳到主要内容

Multi-faceted contrastive learning with inter-frame difference for traffic video question answering

  • Kan Guo*
  • , Qi Zuo
  • , Yongli Hu
  • , Lanping Qian
  • , Daxin Tian
  • , Jiapu Wang
  • , Guixian Qu
  • , Tingzheng Jia
  • , Junbin Gao
  • , Baocai Yin
  • , Jian Wu
  • *此作品的通讯作者
  • Beijing University of Technology
  • The University of Sydney
  • DeepGlint

科研成果: 期刊稿件文章同行评审

摘要

Most existing Video Question Answering (VQA) models rely on vision-language pretraining to bridge the modality gap. However, they often underperform on Traffic Video Question Answering (TrafficVQA) due to its distinctive spatio-temporal characteristics. Surveillance videos often feature static backgrounds with moving objects, while in-car videos involve dynamic camera motion and emphasize interactions among traffic participants. More critically, traffic accidents are usually characterized by sudden and short-lived motion patterns–such as abrupt braking or collisions–that are difficult to capture using models primarily designed for long-term dependencies. To address these challenges, we propose MCL-ID, a Multi-faceted Contrastive Learning framework with Inter-frame Difference analysis. Specifically, we design an inter-frame difference module on top of an image-pretrained backbone to highlight abrupt motion by computing pixel-wise frame differences. We further introduce a cross-attention fusion gate to align spatially localized motion cues with visual features under question guidance. Finally, we incorporate a multi-faceted contrastive learning strategy to enhance cross-modal alignment among motion differences, visual content, and textual queries. Experiments on SUTD-TrafficQA and NExT-QA demonstrate that our approach achieves superior performance, validating its effectiveness for traffic-centric VQA tasks. Code is available at https://github.com/nmjhg/MCL-ID.

源语言英语
文章编号115852
期刊Knowledge-Based Systems
341
DOI
出版状态已出版 - 23 5月 2026

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 3 - 良好健康与福祉
    可持续发展目标 3 良好健康与福祉

学术指纹

探究 'Multi-faceted contrastive learning with inter-frame difference for traffic video question answering' 的科研主题。它们共同构成独一无二的学术指纹。

引用此