Skip to main navigation Skip to search Skip to main content

针对音频识别的物理世界音素对抗攻击

Translated title of the contribution: Phonemic Adversarial Attack Against Audio Recognition in Physical World
  • Jiakai Wang
  • , Yusheng Kong
  • , Zhendong Chen
  • , Jin Hu
  • , Zixin Yin
  • , Yuqing Ma
  • , Qinghong Yang
  • , Xianglong Liu*
  • *Corresponding author for this work
  • Zhongguancun Laboratory
  • Beihang University
  • Polixir Technologies
  • Hefei Comprehensive National Science Center

Research output: Contribution to journalArticlepeer-review

Abstract

Audio recognition has been widely applied in the typical scenarios, like auto-driving, Internet of things, and etc. In recent years, research on adversarial attacks in audio recognition has attracted extensive attention. However, most of the existing studies mainly rely on the coarse-grain audio features at the instance level, which leads to expensive generation time costs and weak universal attacking ability in real world. To address the problem, we propose a phonemic adversarial noise (PAN) generation paradigm, which exploits the audio features at the phoneme level to perform fast and universal adversarial attacks. Experiments are conducted using a variety of datasets commonly used in speech recognition tasks, such as LibriSpeech, to experimentally validate the effectiveness of the PAN proposed in this paper, its ability to generalize across datasets, its ability to migrate attacks across models, and its ability to migrate attacks across tasks, as well as further validating the effectiveness of the attack civilian-oriented Internet of things audio recognition application in the physical world devices. Extensive experiments demonstrate that the proposed PAN outperforms the comparative baselines by large margins (about 24 times speedup and 38% attacking ability improvement on average), and the sampling strategy and learning method proposed in this paper are significant in reducing the training time and improving the attack capability.

Translated title of the contributionPhonemic Adversarial Attack Against Audio Recognition in Physical World
Original languageChinese (Traditional)
Pages (from-to)751-764
Number of pages14
JournalJisuanji Yanjiu yu Fazhan/Computer Research and Development
Volume62
Issue number3
DOIs
StatePublished - Mar 2025

Fingerprint

Dive into the research topics of 'Phonemic Adversarial Attack Against Audio Recognition in Physical World'. Together they form a unique fingerprint.

Cite this