Skip to main navigation Skip to search Skip to main content

Cross-Modal Contrastive Masked AutoEncoder for Compressed Video Pre-Training

  • Bing Li
  • , Jiaxin Chen*
  • , Guohao Li
  • , Dongming Zhang
  • , Xiuguo Bao
  • , Di Huang
  • *Corresponding author for this work
  • Beihang University
  • National Computer Network Emergency Response Technical Team

Research output: Contribution to journalArticlepeer-review

Abstract

In this paper, we propose a novel Transformer based approach, namely Cross-modal Contrastive Masked AutoEncoder (C2MAE), to Self-Supervised Learning (SSL) on compressed videos. A unified Transformer encoder is employed to discover relationships of visual tokens from RGBs, motion vectors and residuals. A hybrid SSL framework is proposed, which combines the complementary advantages of Masked Image Modeling (MIM) and Contrastive Learning (CL) pretext tasks, for powerful representation learning. The MIM branch extends VideoMAE by a new Fine-Grained Motion-aware Masking (FGMM) strategy and a modified Multi-modal Reconstruction (MR) task, where FGMM computes motion saliency maps as motion priors to guide the masks so that it well fits for the data properties in the compressed domain and the MR task highlights the reconstruction of raw videos by joint representations from corresponding compressed videos in addition to that in each single modality. The CL branch introduces the Contrastive Cross-modal Learning (CCL) module, and the features from a compressed video clip and the ones from its raw video counterpart are compared instead of widely used augmented data. Due to these designs, C2MAE significantly enhances interactions across modalities to compensate the sparsity of I-frames and the coarse and noisy nature of P-frames, thus delivering much stronger pre-trained models. Extensive experiments are conducted on the UCF-101, HMDB-51 and Kinetics-400 benchmarks with state-of-the-art results reported, demonstrating its effectiveness.

Original languageEnglish
Pages (from-to)4500-4514
Number of pages15
JournalIEEE Transactions on Image Processing
Volume34
DOIs
StatePublished - 2025

Keywords

  • Compressed video
  • action recognition
  • contrastive learning
  • masked image modeling
  • video retrieval

Fingerprint

Dive into the research topics of 'Cross-Modal Contrastive Masked AutoEncoder for Compressed Video Pre-Training'. Together they form a unique fingerprint.

Cite this