Skip to main navigation Skip to search Skip to main content

Bullet-Screen-Emoji Attack With Temporal Difference Noise for Video Action Recognition

  • Yongkang Zhang
  • , Han Zhang
  • , Jun Li*
  • , Zhiping Shi
  • , Jian Yang
  • , Kaixin Yang
  • , Shuo Yin
  • , Qiuyan Liang
  • , Xianglong Liu
  • *Corresponding author for this work
  • Capital Normal University
  • Beihang University
  • University of Chinese Academy of Sciences

Research output: Contribution to journalArticlepeer-review

Abstract

Recent studies have shown that video action recognition models are also vulnerable to fooling by adversarial samples. However, currently existing video attack methods usually require high computational overhead (e.g., they generate adversarial perturbations for all frames by default), and most of them are difficult to implement printable attacks in the physical world. To address the above issues, we devise a novel efficient and effective framework for video action recognition attack: Bullet-Screen-Emoji Attack with Temporal Difference Noise (BSE), a reinforcement learning-based black-box attack method that fools the model by simply generating adversarial bullet screens for key frame and scrolling them on clean video. The agent is optimized to make the optimal actions, i.e., searching key frame. Moreover, we introduce a simple and effective temporal difference noise to enhance the attack capability of the adversarial bullet screen and accelerate the convergence speed. Most importantly, BSE enables printable physical attacks. Extensive experiments show that our proposed BSE achieves promising attack performance on mainstream datasets (HMDB51, UCF101 and Kinetics-400) and in the physical world with high efficiency.

Original languageEnglish
Pages (from-to)589-600
Number of pages12
JournalIEEE Transactions on Circuits and Systems for Video Technology
Volume35
Issue number1
DOIs
StatePublished - 2025

Keywords

  • Video adversarial attack
  • reinforcement learning
  • temporal difference noise

Fingerprint

Dive into the research topics of 'Bullet-Screen-Emoji Attack With Temporal Difference Noise for Video Action Recognition'. Together they form a unique fingerprint.

Cite this