Skip to main navigation Skip to search Skip to main content

AST-Adapter: Parameter-Efficient Video-to-Video Transfer Learning With Adaptive Spatiotemporal Information Bias

  • Puyue Hou
  • , Guohao Li
  • , Zhi Cai
  • , Jinjin Zhang
  • , Di Huang*
  • *Corresponding author for this work
  • Beihang University

Research output: Contribution to journalArticlepeer-review

Abstract

Leveraging video pre-trained models for video downstream tasks has recently emerged with promising performance. Except for the full fine-tuning paradigm, parameter-efficient transfer learning (PETL) exists as a promising way and has not yet been fully explored in video-to-video transfer learning. While current PETL approaches succeed to reduce parameter quantity and computation cost, they overlook the critical spatiotemporal property in video modality. In this paper, we first propose a novel metric to quantify the spatiotemporal information bias across video datasets and uncover its impact on transfer pReferences through systematic analysis. Based on the analysis, we introduce an innovative parameter-efficient transfer learning method, named Adaptive SpatioTemporal Adapter (AST-Adapter). Our approach automatically adjusts layer-wise architectures with different spatiotemporal adapter modules to exploit the intrinsics of downstream tasks to achieve adaptive spatiotemporal learning, thus delivering robustness and generalization. Extensive experiments on five datasets across action recognition and action detection task show that AST-Adapter surpasses both video-to-video and image-to-video approaches, whilst keeping the advantage of parameter efficiency. Notably, AST-Adapter achieves 89.9% on Kinectics-400, 80.3% on HMDB51 and 97.7% on UCF101 while introduces only 1% to 10% tunable parameters. Our code is available at https://github.com/hhhhhpy/AST-Adapter

Original languageEnglish
Pages (from-to)4858-4872
Number of pages15
JournalIEEE Transactions on Circuits and Systems for Video Technology
Volume36
Issue number4
DOIs
StatePublished - 2026

Keywords

  • Video pre-trained model
  • parameter-efficient transfer learning
  • spatiotemporal information bias

Fingerprint

Dive into the research topics of 'AST-Adapter: Parameter-Efficient Video-to-Video Transfer Learning With Adaptive Spatiotemporal Information Bias'. Together they form a unique fingerprint.

Cite this