Page 363 - 《软件学报》2026年第7期
P. 363
3048 软件学报 2026 年第 37 卷第 7 期
[40] Rix AW, Beerends JG, Hollier MP, Hekstra AP. Perceptual evaluation of speech quality (PESQ)—A new method for speech quality
assessment of telephone networks and codecs. In: Proc. of the 2001 IEEE Int’l Conf. on Acoustics, Speech, and Signal Processing. Salt
Lake City: IEEE, 2001. 749–752. [doi: 10.1109/ICASSP.2001.941023]
[41] Goodfellow IJ, Shlens J, Szegedy C. Explaining and harnessing adversarial examples. In: Proc. of the 3rd Int’l Conf. on Learning
Representations. San Diego: OpenReview.net, 2015.
[42] Mądry A, Makelov A, Schmidt L, Tsipras D, Vladu A. Towards deep learning models resistant to adversarial attacks. In: Proc. of the 6th
Int’l Conf. on Learning Representations. Vancouver: OpenReview.net, 2018. 1050.
[43] Carlini N, Wagner D. Audio adversarial examples: Targeted attacks on speech-to-text. In: Proc. of the 2018 IEEE Security and Privacy
Workshops. San Francisco: IEEE, 2018. 1–7. [doi: 10.1109/SPW.2018.00009]
[44] Du TY, Ji SL, Li JF, Gu QC, Wang T, Beyah R. SirenAttack: Generating adversarial audio for end-to-end acoustic systems. In: Proc. of
the 15th ACM Asia Conf. on Computer and Communications Security. Taipei: ACM, 2020. 357–369. [doi: 10.1145/3320269.3384733]
[45] Abdullah H, Rahman MS, Garcia W, Warren K, Yadav AS, Shrimpton T, Traynor P. Hear “no evil”, see “kenansville”: Efficient and
transferable black-box attacks on speech recognition and voice identification systems. In: Proc. of the 2021 IEEE Symp. on Security and
Privacy. San Francisco: IEEE, 2021. 712–729. [doi: 10.1109/SP40001.2021.00009]
[46] Yu ZY, Chang Y, Zhang N, Xiao CW. SMACK: Semantically meaningful adversarial audio attack. In: Proc. of the 32nd USENIX
Security Symp. Anaheim: USENIX Association, 2023. 3799–3816.
[47] Wang XB, Hou R, Zhao BY, Yuan FK, Zhang J, Meng D, Qian XH. DNNGuard: An elastic heterogeneous DNN accelerator architecture
against adversarial attacks. In: Proc. of the 25th Int’l Conf. on Architectural Support for Programming Languages and Operating Systems.
Lausanne: ACM, 2020. 19–34. [doi: 10.1145/3373376.3378532]
[48] Credamo. 2025 (in Chinese). https://www.credamo.com/home.html#/
[49] Joshi S, Villalba J, Żelasko P, Moro-Velázquez L, Dehak N. Study of pre-processing defenses against adversarial attacks on state-of-the-
art speaker recognition systems. IEEE Trans. on Information Forensics and Security, 2021, 16: 4811–4826. [doi: 10.1109/TIFS.2021.
3116438]
附中文参考文献
[37] 腾讯. 腾讯云语音识别. 2025. https://cloud.tencent.com/product/asr
[38] 讯飞. 声纹识别. 2025. https://www.xfyun.cn/services/voiceprint-recognition
[39] 云知声. 声纹识别. 2025. https://ai-poc.hivoice.cn/voiceprint-recognition
[48] 见数. 2025. https://www.credamo.com/home.html#/
作者简介
陈家源, 硕士生, 主要研究领域为语音对抗攻击和防御.
黄文弘, 博士生, 主要研究领域为人工智能安全.
黄方军, 博士, 教授, 博士生导师, CCF 高级会员, 主要研究领域为人工智能安全, 多媒体内容安全.

