Page 362 - 《软件学报》2026年第7期
P. 362
陈家源 等: 说话人信息引导的高性能音频对抗攻击 3047
perturbations. In: Proc. of the 2020 ACM SIGSAC Conf. on Computer and Communications Security. ACM, 2020. 1121–1134. [doi: 10.
1145/3372297.3423348]
[17] Lin YQ, Abdulla WH. Principles of psychoacoustics. In: Lin YQ, Abdulla WH, eds. Audio Watermark: A Comprehensive Foundation
Using Matlab. Cham: Springer, 2015. 15–49. [doi: 10.1007/978-3-319-07974-5_2]
[18] Xie CH, Zhang ZS, Zhou YY, Bai S, Wang JY, Ren Z, Yuille AL. Improving transferability of adversarial examples with input diversity.
In: Proc. of the 2019 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 2725–2734. [doi: 10.1109/
CVPR.2019.00284]
[19] Chen GK, Zhang YD, Zhao Z, Song F. QFA2SR: Query-free adversarial transfer attacks to speaker recognition systems. In: Proc. of the
32nd USENIX Security Symp. Anaheim: USENIX Association, 2023. 2437–2454.
[20] Zhang Y, Li HW, Xu GW, Luo XZ, Dong GS. Generating audio adversarial examples with ensemble substituted models. In: Proc. of the
2021 IEEE Int’l Conf. on Communications. Montreal: IEEE, 2021. 1–6. [doi: 10.1109/ICC42927.2021.9500431]
[21] Xue M, Peng K, Gong XL, Zhang Q, Chen YJ, Li RT. Echo: Reverberation-based fast black-box adversarial attacks on intelligent audio
systems. Proc. of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2023, 7(3): 137. [doi: 10.1145/3610874]
[22] Reynolds DA, Quatieri TF, Dunn RB. Speaker verification using adapted Gaussian mixture models. Digital Signal Processing, 2000,
10(1–3): 19–41. [doi: 10.1006/dspr.1999.0361]
[23] Dehak N, Dehak R, Kenny P, Brümmer N, Ouellet P, Dumouchel P. Support vector machines versus fast scoring in the low-dimensional
total variability space for speaker verification. In: Proc. of the 10th Annual Conf. of the Int’l Speech Communication Association.
Brighton: ISCA, 2009. 1559–1562.
[24] Variani E, Lei X, McDermott E, Moreno IL, Gonzalez-Dominguez J. Deep neural networks for small footprint text-dependent speaker
verification. In: Proc. of the 2014 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing. Florence: IEEE, 2014. 4052–4056. [doi:
10.1109/ICASSP.2014.6854363]
[25] Snyder D, Garcia-Romero D, Sell G, Povey D, Khudanpur S. X-vectors: Robust DNN embeddings for speaker recognition. In: Proc. of
the 2018 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing. Calgary: IEEE, 2018. 5329–5333. [doi: 10.1109/ICASSP.2018.
8461375]
[26] Chen GK, Zhao Z, Song F, Chen S, Fan LL, Liu Y. AS2T: Arbitrary source-to-target adversarial attack on speaker recognition systems.
IEEE Trans. on Dependable and Secure Computing, 2022: 1–17. [doi: 10.1109/TDSC.2022.3189397]
[27] Li JY, Tu WP, Xiao L. FreeVC: Towards high-quality text-free one-shot voice conversion. In: Proc. of the 2023 IEEE Int’l Conf. on
Acoustics, Speech and Signal Processing. Rhodes Island: IEEE, 2023. 1–5. [doi: 10.1109/ICASSP49357.2023.10095191]
[28] Liu SX, Cao YW, Wang DS, Wu XX, Liu XY, Meng HL. Any-to-many voice conversion with location-relative sequence-to-sequence
modeling. IEEE/ACM Trans. on Audio, Speech, and Language Processing, 2021, 29: 1717–1728. [doi: 10.1109/TASLP.2021.3076867]
[29] Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In: Proc. of the 38th Int’l
Conf. on Machine Learning. PMLR, 2021. 5530–5540.
[30] Panayotov V, Chen GG, Povey D, Khudanpur S. LibriSpeech: An ASR corpus based on public domain audio books. In: Proc. of the 2015
IEEE Int’l Conf. on Acoustics, Speech and Signal Processing. South Brisbane: IEEE, 2015. 5206–5210. [doi: 10.1109/ICASSP.2015.
7178964]
[31] Wu ZZ, Khodabakhsh A, Demiroglu C, Yamagishi J, Saito D, Toda T, King S. SAS: A speaker verification spoofing database containing
diverse attacks. In: Proc. of the 2015 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing. South Brisbane: IEEE, 2015.
4440–4444. [doi: 10.1109/ICASSP.2015.7178810]
[32] Chen GK, Zhao Z, Song F, Chen S, Fan LL, Wang F, Wang J. Towards understanding and mitigating audio adversarial examples for
speaker recognition. IEEE Trans. on Dependable and Secure Computing, 2023, 20(5): 3970–3987. [doi: 10.1109/TDSC.2022.3220673]
[33] KALDI. VoxCeleb Models. 2022. https://kaldi-asr.org/models/m7
[34] KALDI. SITW Models. 2022. https://kaldi-asr.org/models/m8
[35] Desplanques B, Thienpondt J, Demuynck K. ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN
based speaker verification. In: Proc. of the 21st Annual Conf. of the Int’l Speech Communication Association. Shanghai: ISCA, 2020.
3830–3834.
[36] Faundez-Zanuy M, Monte-Moreno E. State-of-the-art in speaker recognition. IEEE Aerospace and Electronic Systems Magazine, 2005,
20(5): 7–12. [doi: 10.1109/MAES.2005.1432568]
[37] Tencent Cloud. Tencent automatic speech recognition. 2025 (in Chinese). https://cloud.tencent.com/product/asr
[38] IFlytek. IFlytek voiceprint recognition. 2025 (in Chinese). https://www.xfyun.cn/services/voiceprint-recognition
[39] Unisound. Unisound voiceprint recognition. 2025 (in Chinese). https://ai-poc.hivoice.cn/voiceprint-recognition

