摘要
Most existing Video Question Answering (VQA) models rely on vision-language pretraining to bridge the modality gap. However, they often underperform on Traffic Video Question Answering (TrafficVQA) due to its distinctive spatio-temporal characteristics. Surveillance videos often feature static backgrounds with moving objects, while in-car videos involve dynamic camera motion and emphasize interactions among traffic participants. More critically, traffic accidents are usually characterized by sudden and short-lived motion patterns–such as abrupt braking or collisions–that are difficult to capture using models primarily designed for long-term dependencies. To address these challenges, we propose MCL-ID, a Multi-faceted Contrastive Learning framework with Inter-frame Difference analysis. Specifically, we design an inter-frame difference module on top of an image-pretrained backbone to highlight abrupt motion by computing pixel-wise frame differences. We further introduce a cross-attention fusion gate to align spatially localized motion cues with visual features under question guidance. Finally, we incorporate a multi-faceted contrastive learning strategy to enhance cross-modal alignment among motion differences, visual content, and textual queries. Experiments on SUTD-TrafficQA and NExT-QA demonstrate that our approach achieves superior performance, validating its effectiveness for traffic-centric VQA tasks. Code is available at https://github.com/nmjhg/MCL-ID.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 115852 |
| 期刊 | Knowledge-Based Systems |
| 卷 | 341 |
| DOI | |
| 出版状态 | 已出版 - 23 5月 2026 |
联合国可持续发展目标
此成果有助于实现下列可持续发展目标:
-
可持续发展目标 3 良好健康与福祉
学术指纹
探究 'Multi-faceted contrastive learning with inter-frame difference for traffic video question answering' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver