Page 361 - 《软件学报》2026年第7期
P. 361
3046 软件学报 2026 年第 37 卷第 7 期
由图 4 可知, 不同模型所需的迭代步数存在差异. 具体而言, 针对 IV_PLDA 和 XV_PLDA 模型发起攻击的难
度相对较低, 其所需的迭代步数较少; 而针对 ResNet 和 ECAPA 模型的攻击难度则相对较高, 需要更多的迭代次
数才能达到预期效果. 此外, 不同置信度水平下的攻击效率呈现出明显的差别. 置信度越高, 需要进行的迭代次数
κ = 3/0.1 的设置情形下, 平均需要 300 次迭代才能够完
就越多, 同时生成的对抗音频的攻击成功率也相应越高. 在
成攻击. 虽然提高置信度在一定程度上能够提升攻击效果, 但这也会导致所需的迭代次数增多, 进而使得资源消耗
大幅增加.
5 总 结
本文提出了一种基于说话人信息的音频对抗攻击方法——SIAttack. 该方法首先对原始音频进行信息解耦, 随
后在解耦得到的说话人信息上通过迭代优化施加合理范围内的对抗扰动, 并重构生成对抗音频, 以有效攻击说话
人识别系统. 大量实验结果表明, SIAttack 生成的对抗音频能够成功误导当前主流说话人识别模型. 与现有对抗攻
击方法相比, SIAttack 所生成的样本在迁移性、隐蔽性与鲁棒性这 3 个方面均表现出更优的性能. 本研究揭示了
在音频深层语义信息层面实施对抗攻击的可行性, 为音频对抗攻击领域的研究提供了新的方向.
References
[1] Kinnunen T, Li HZ. An overview of text-independent speaker recognition: From features to supervectors. Speech Communication, 2010,
52(1): 12–40. [doi: 10.1016/j.specom.2009.08.009]
[2] Markowitz JA. Voice biometrics. Communications of the ACM, 2000, 43(9): 66–73. [doi: 10.1145/348941.348995]
[3] Singh S. Forensic and automatic speaker recognition system. Int’l Journal of Electrical and Computer Engineering (IJECE), 2018, 8(5):
2804–2811. [doi: 10.11591/ijece.v8i5.pp2804-2811]
[4] Gu ZQ, Hu WX, Zhang CJ, Lu H, Yin LH, Wang L. Gradient shielding: Towards understanding vulnerability of deep neural networks.
IEEE Trans. on Network Science and Engineering, 2021, 8(2): 921–932. [doi: 10.1109/TNSE.2020.2996738]
[5] Huang WH, Dai YS, Fei JW, Huang FJ. New visible watermark protection mechanism based on information hiding. IEEE Trans. on
Information Forensics and Security, 2025, 20: 7764–7776. [doi: 10.1109/TIFS.2025.3592572]
[6] Lu H, Jin CJ, Helu XH, Zhu CS, Guizani N, Tian ZH. AutoD: Intelligent blockchain application unpacking based on JNI layer deception
call. IEEE Network, 2021, 35(2): 215–221. [doi: 10.1109/MNET.011.2000467]
[7] Kreuk F, Adi Y, Cisse M, Keshet J. Fooling end-to-end speaker verification with adversarial examples. In: Proc. of the 2018 IEEE Int’l
Conf. on Acoustics, Speech and Signal Processing. Calgary: IEEE, 2018. 1962–1966. [doi: 10.1109/ICASSP.2018.8462693]
[8] Li ZH, Shi C, Xie Y, Liu J, Yuan B, Chen YY. Practical adversarial attacks against speaker recognition systems. In: Proc. of the 21st Int’l
Workshop on Mobile Computing Systems and Applications. Austin: ACM, 2020. 9–14. [doi: 10.1145/3376897.3377856]
[9] Xie Y, Shi C, Li ZH, Liu J, Chen YY, Yuan B. Real-time, universal, and robust adversarial attacks against speaker recognition systems.
In: Proc. of the 2020 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing. Barcelona: IEEE, 2020. 1738–1742. [doi: 10.1109/
ICASSP40776.2020.9053747]
[10] Chen GK, Chenb S, Fan LL, Du XN, Zhao Z, Song F, Liu Y. Who is real Bob? Adversarial attacks on speaker recognition systems. In:
Proc. of the 2021 IEEE Symp. on Security and Privacy. San Francisco: IEEE, 2021. 694–711. [doi: 10.1109/SP40001.2021.00004]
[11] Qin Y, Carlini N, Cottrell G, Goodfellow I, Raffel C. Imperceptible, robust, and targeted adversarial examples for automatic speech
recognition. In: Proc. of the 36th Int’l Conf. on Machine Learning. Long Beach: PMLR, 2019. 5231–5240.
[12] Wang Q, Guo PC, Xie L. Inaudible adversarial perturbations for targeted attack in speaker recognition. In: Proc. of the 21st Annual Conf.
of the Int’l Speech Communication Association. Shanghai: ISCA, 2020. 4228–4232.
[13] Zhang GM, Yan C, Ji XY, Zhang TC, Zhang TM, Xu WY. DolphinAttack: Inaudible voice commands. In: Proc. of the 2017 ACM
SIGSAC Conf. on Computer and Communications Security. Dallas: ACM, 2017. 103–117. [doi: 10.1145/3133956.3134052]
[14] Chen T, Shangguan LF, Li ZJ, Jamieson K. Metamorph: Injecting inaudible commands into over-the-air voice controlled systems. In:
Proc. of the 27th Annual Network and Distributed System Security Symp. San Diego: The Internet Society, 2020.
[15] Yuan XJ, Chen YX, Zhao Y, Long YH, Liu XK, Chen K, Zhang SZ, Huang HQ, Wang XF, Gunter CA. CommanderSong: A systematic
approach for practical adversarial voice recognition. In: Proc. of the 27th USENIX Security Symp. Baltimore: USENIX Association,
2018. 49–64.
[16] Li ZH, Wu Y, Liu J, Chen YY, Yuan B. AdvPulse: Universal, synchronization-free, and targeted audio adversarial attacks via subsecond

