跳到主要导航 跳到搜索 跳到主要内容

Bullet-Screen-Emoji Attack With Temporal Difference Noise for Video Action Recognition

  • Yongkang Zhang
  • , Han Zhang
  • , Jun Li*
  • , Zhiping Shi
  • , Jian Yang
  • , Kaixin Yang
  • , Shuo Yin
  • , Qiuyan Liang
  • , Xianglong Liu
  • *此作品的通讯作者
  • Capital Normal University
  • Beihang University
  • University of Chinese Academy of Sciences

科研成果: 期刊稿件文章同行评审

摘要

Recent studies have shown that video action recognition models are also vulnerable to fooling by adversarial samples. However, currently existing video attack methods usually require high computational overhead (e.g., they generate adversarial perturbations for all frames by default), and most of them are difficult to implement printable attacks in the physical world. To address the above issues, we devise a novel efficient and effective framework for video action recognition attack: Bullet-Screen-Emoji Attack with Temporal Difference Noise (BSE), a reinforcement learning-based black-box attack method that fools the model by simply generating adversarial bullet screens for key frame and scrolling them on clean video. The agent is optimized to make the optimal actions, i.e., searching key frame. Moreover, we introduce a simple and effective temporal difference noise to enhance the attack capability of the adversarial bullet screen and accelerate the convergence speed. Most importantly, BSE enables printable physical attacks. Extensive experiments show that our proposed BSE achieves promising attack performance on mainstream datasets (HMDB51, UCF101 and Kinetics-400) and in the physical world with high efficiency.

源语言英语
页(从-至)589-600
页数12
期刊IEEE Transactions on Circuits and Systems for Video Technology
35
1
DOI
出版状态已出版 - 2025

学术指纹

探究 'Bullet-Screen-Emoji Attack With Temporal Difference Noise for Video Action Recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此