Abstract
Most existing Video Question Answering (VQA) models rely on vision-language pretraining to bridge the modality gap. However, they often underperform on Traffic Video Question Answering (TrafficVQA) due to its distinctive spatio-temporal characteristics. Surveillance videos often feature static backgrounds with moving objects, while in-car videos involve dynamic camera motion and emphasize interactions among traffic participants. More critically, traffic accidents are usually characterized by sudden and short-lived motion patterns–such as abrupt braking or collisions–that are difficult to capture using models primarily designed for long-term dependencies. To address these challenges, we propose MCL-ID, a Multi-faceted Contrastive Learning framework with Inter-frame Difference analysis. Specifically, we design an inter-frame difference module on top of an image-pretrained backbone to highlight abrupt motion by computing pixel-wise frame differences. We further introduce a cross-attention fusion gate to align spatially localized motion cues with visual features under question guidance. Finally, we incorporate a multi-faceted contrastive learning strategy to enhance cross-modal alignment among motion differences, visual content, and textual queries. Experiments on SUTD-TrafficQA and NExT-QA demonstrate that our approach achieves superior performance, validating its effectiveness for traffic-centric VQA tasks. Code is available at https://github.com/nmjhg/MCL-ID.
| Original language | English |
|---|---|
| Article number | 115852 |
| Journal | Knowledge-Based Systems |
| Volume | 341 |
| DOIs | |
| State | Published - 23 May 2026 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- Contrastive learning
- Inter-frame difference
- TrafficVQA
Fingerprint
Dive into the research topics of 'Multi-faceted contrastive learning with inter-frame difference for traffic video question answering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver