Skip to main navigation Skip to search Skip to main content

Multi-faceted contrastive learning with inter-frame difference for traffic video question answering

  • Kan Guo*
  • , Qi Zuo
  • , Yongli Hu
  • , Lanping Qian
  • , Daxin Tian
  • , Jiapu Wang
  • , Guixian Qu
  • , Tingzheng Jia
  • , Junbin Gao
  • , Baocai Yin
  • , Jian Wu
  • *Corresponding author for this work
  • Beijing University of Technology
  • The University of Sydney
  • DeepGlint

Research output: Contribution to journalArticlepeer-review

Abstract

Most existing Video Question Answering (VQA) models rely on vision-language pretraining to bridge the modality gap. However, they often underperform on Traffic Video Question Answering (TrafficVQA) due to its distinctive spatio-temporal characteristics. Surveillance videos often feature static backgrounds with moving objects, while in-car videos involve dynamic camera motion and emphasize interactions among traffic participants. More critically, traffic accidents are usually characterized by sudden and short-lived motion patterns–such as abrupt braking or collisions–that are difficult to capture using models primarily designed for long-term dependencies. To address these challenges, we propose MCL-ID, a Multi-faceted Contrastive Learning framework with Inter-frame Difference analysis. Specifically, we design an inter-frame difference module on top of an image-pretrained backbone to highlight abrupt motion by computing pixel-wise frame differences. We further introduce a cross-attention fusion gate to align spatially localized motion cues with visual features under question guidance. Finally, we incorporate a multi-faceted contrastive learning strategy to enhance cross-modal alignment among motion differences, visual content, and textual queries. Experiments on SUTD-TrafficQA and NExT-QA demonstrate that our approach achieves superior performance, validating its effectiveness for traffic-centric VQA tasks. Code is available at https://github.com/nmjhg/MCL-ID.

Original languageEnglish
Article number115852
JournalKnowledge-Based Systems
Volume341
DOIs
StatePublished - 23 May 2026

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Contrastive learning
  • Inter-frame difference
  • TrafficVQA

Fingerprint

Dive into the research topics of 'Multi-faceted contrastive learning with inter-frame difference for traffic video question answering'. Together they form a unique fingerprint.

Cite this