Page 363 - 《软件学报》2026年第7期
P. 363

3048                                                       软件学报  2026  年第  37  卷第  7  期


                 [40]   Rix AW, Beerends JG, Hollier MP, Hekstra AP. Perceptual evaluation of speech quality (PESQ)—A new method for speech quality
                     assessment of telephone networks and codecs. In: Proc. of the 2001 IEEE Int’l Conf. on Acoustics, Speech, and Signal Processing. Salt
                     Lake City: IEEE, 2001. 749–752. [doi: 10.1109/ICASSP.2001.941023]
                 [41]   Goodfellow  IJ,  Shlens  J,  Szegedy  C.  Explaining  and  harnessing  adversarial  examples.  In:  Proc.  of  the  3rd  Int’l  Conf.  on  Learning
                     Representations. San Diego: OpenReview.net, 2015.
                 [42]   Mądry A, Makelov A, Schmidt L, Tsipras D, Vladu A. Towards deep learning models resistant to adversarial attacks. In: Proc. of the 6th
                     Int’l Conf. on Learning Representations. Vancouver: OpenReview.net, 2018. 1050.
                 [43]   Carlini N, Wagner D. Audio adversarial examples: Targeted attacks on speech-to-text. In: Proc. of the 2018 IEEE Security and Privacy
                     Workshops. San Francisco: IEEE, 2018. 1–7. [doi: 10.1109/SPW.2018.00009]
                 [44]   Du TY, Ji SL, Li JF, Gu QC, Wang T, Beyah R. SirenAttack: Generating adversarial audio for end-to-end acoustic systems. In: Proc. of
                     the 15th ACM Asia Conf. on Computer and Communications Security. Taipei: ACM, 2020. 357–369. [doi: 10.1145/3320269.3384733]
                 [45]   Abdullah H, Rahman MS, Garcia W, Warren K, Yadav AS, Shrimpton T, Traynor P. Hear “no evil”, see “kenansville”: Efficient and
                     transferable black-box attacks on speech recognition and voice identification systems. In: Proc. of the 2021 IEEE Symp. on Security and
                     Privacy. San Francisco: IEEE, 2021. 712–729. [doi: 10.1109/SP40001.2021.00009]
                 [46]   Yu  ZY,  Chang  Y,  Zhang  N,  Xiao  CW.  SMACK:  Semantically  meaningful  adversarial  audio  attack.  In:  Proc.  of  the  32nd  USENIX
                     Security Symp. Anaheim: USENIX Association, 2023. 3799–3816.
                 [47]   Wang XB, Hou R, Zhao BY, Yuan FK, Zhang J, Meng D, Qian XH. DNNGuard: An elastic heterogeneous DNN accelerator architecture
                     against adversarial attacks. In: Proc. of the 25th Int’l Conf. on Architectural Support for Programming Languages and Operating Systems.
                     Lausanne: ACM, 2020. 19–34. [doi: 10.1145/3373376.3378532]
                 [48]   Credamo. 2025 (in Chinese). https://www.credamo.com/home.html#/
                 [49]   Joshi S, Villalba J, Żelasko P, Moro-Velázquez L, Dehak N. Study of pre-processing defenses against adversarial attacks on state-of-the-
                     art speaker recognition systems. IEEE Trans. on Information Forensics and Security, 2021, 16: 4811–4826. [doi: 10.1109/TIFS.2021.
                     3116438]

                 附中文参考文献
                 [37]   腾讯. 腾讯云语音识别. 2025. https://cloud.tencent.com/product/asr
                 [38]   讯飞. 声纹识别. 2025. https://www.xfyun.cn/services/voiceprint-recognition
                 [39]   云知声. 声纹识别. 2025. https://ai-poc.hivoice.cn/voiceprint-recognition
                 [48]   见数. 2025. https://www.credamo.com/home.html#/

                 作者简介
                 陈家源, 硕士生, 主要研究领域为语音对抗攻击和防御.
                 黄文弘, 博士生, 主要研究领域为人工智能安全.
                 黄方军, 博士, 教授, 博士生导师, CCF  高级会员, 主要研究领域为人工智能安全, 多媒体内容安全.
   358   359   360   361   362   363   364