摘要
Audio recognition has been widely applied in the typical scenarios, like auto-driving, Internet of things, and etc. In recent years, research on adversarial attacks in audio recognition has attracted extensive attention. However, most of the existing studies mainly rely on the coarse-grain audio features at the instance level, which leads to expensive generation time costs and weak universal attacking ability in real world. To address the problem, we propose a phonemic adversarial noise (PAN) generation paradigm, which exploits the audio features at the phoneme level to perform fast and universal adversarial attacks. Experiments are conducted using a variety of datasets commonly used in speech recognition tasks, such as LibriSpeech, to experimentally validate the effectiveness of the PAN proposed in this paper, its ability to generalize across datasets, its ability to migrate attacks across models, and its ability to migrate attacks across tasks, as well as further validating the effectiveness of the attack civilian-oriented Internet of things audio recognition application in the physical world devices. Extensive experiments demonstrate that the proposed PAN outperforms the comparative baselines by large margins (about 24 times speedup and 38% attacking ability improvement on average), and the sampling strategy and learning method proposed in this paper are significant in reducing the training time and improving the attack capability.
| 投稿的翻译标题 | Phonemic Adversarial Attack Against Audio Recognition in Physical World |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 751-764 |
| 页数 | 14 |
| 期刊 | Jisuanji Yanjiu yu Fazhan/Computer Research and Development |
| 卷 | 62 |
| 期 | 3 |
| DOI | |
| 出版状态 | 已出版 - 3月 2025 |
关键词
- adversarial attack
- audio recognition
- phonemic adversarial noise
- physical attack
- security of artificial intelligence
学术指纹
探究 '针对音频识别的物理世界音素对抗攻击' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver