Page 348 - 《软件学报》2026年第7期
P. 348
软件学报 ISSN 1000-9825, CODEN RUXUEW E-mail: jos@iscas.ac.cn
2026,37(7):3033−3048 [doi: 10.13328/j.cnki.jos.007574] [CSTR: 32375.14.jos.007574] http://www.jos.org.cn
©中国科学院软件研究所版权所有. Tel: +86-10-62562563
*
说话人信息引导的高性能音频对抗攻击
陈家源, 黄文弘, 黄方军
(中山大学 网络空间安全学院, 广东 深圳 518107)
通信作者: 黄方军, E-mail: huangfj@mail.sysu.edu.cn
摘 要: 随着音频对抗攻击研究的深入, 如何在确保对抗音频隐蔽性 (即与原始音频在听觉上高度相似) 的同时, 提
高其在不同模型之间的迁移性, 已成为研究热点之一. 提出一种能够同时提高对抗音频隐蔽性和迁移性的方法
SIAttack (speak information attack). 该方法的核心思想是解耦音频中的说话人信息与内容信息, 并仅对说话人信息
施加轻微扰动, 从而可以在保持内容信息不变的前提下实现对说话人识别系统的高效攻击. 在 4 个说话人识别模
型以及 3 个主流商业 API 上的实验表明, SIAttack 生成的音频在听觉上几乎无法与原始音频区分, 且能以较高的
成功率误导所有测试模型, 在说话人识别模型上迁移成功率最高可达 100%.
关键词: 音频; 隐蔽性; 迁移性; 对抗攻击; 对抗扰动
中图法分类号: TP309
中文引用格式: 陈家源, 黄文弘, 黄方军. 说话人信息引导的高性能音频对抗攻击. 软件学报, 2026, 37(7): 3033–3048. http://www.
jos.org.cn/1000-9825/7574.htm
英文引用格式: Chen JY, Huang WH, Huang FJ. High-performance Audio Adversarial Attacks Guided by Speaker Information. Ruan
Jian Xue Bao/Journal of Software, 2026, 37(7): 3033–3048 (in Chinese). http://www.jos.org.cn/1000-9825/7574.htm
High-performance Audio Adversarial Attacks Guided by Speaker Information
CHEN Jia-Yuan, HUANG Wen-Hong, HUANG Fang-Jun
(School of Cyber Science and Technology, Sun Yat-sen University, Shenzhen 518107, China)
Abstract: As the research on audio adversarial attacks advances, improving the transferability of adversarial audio across different models
and ensuring its imperceptibility (that is, highly similar to the original audio in auditory perception) at the same time have become a
research hotspot. This study proposes a new method called speak information attack (SIAttack) that can simultaneously improve the
imperceptibility and transferability of adversarial audio. Specifically, the core idea of this method is to decouple speaker information from
content information in the audio, and then apply small perturbations only to the speaker information, thereby achieving efficient attacks on
the speaker recognition system under the premise of keeping the content information unchanged. The experiments on four speaker
recognition models and three mainstream commercial APIs show that the audio generated by SIAttack is almost indistinguishable from the
original audio, and can mislead all test models with a high success rate. Additionally, the transfer success rate on speaker recognition
models can reach up to 100%.
Key words: audio; imperceptibility; transferability; adversarial attack; adversarial perturbation
说话人识别系统通过提取和比对个人独特的声纹特征, 实现对个人身份的识别或验证. 近年来, 随着人工智能
技术的飞速发展, 基于 AI 的说话人识别系统已广泛应用于智能家居 [1] 、支付交易 [2] 以及伪造检测 [3] 等现实场景,
成为现代身份认证体系中的重要组成部分. 然而, 随着该类系统的普及, 所面临的潜在安全风险也日益突出, 尤其
是对抗攻击对系统可靠性构成的严峻挑战 [4,5] . 攻击者可通过向音频信号中注入人耳难以察觉的对抗扰动, 误导说
话人识别系统产生错误判断, 从而严重威胁系统安全 [6] , 引发隐私泄露、金融诈骗和法律纠纷等严重后果.
近年来, 研究者陆续提出了一系列针对说话人识别系统的对抗攻击方法. 这类方法通常通过对原始音频注入
* 基金项目: 国家自然科学基金 (U2336208); 深圳市科技计划 (JCYJ20250604175534044)
收稿时间: 2024-06-14; 修改时间: 2025-04-29; 采用时间: 2025-11-20; jos 在线出版时间: 2026-02-11
CNKI 网络首发时间: 2026-02-12

